Skip to main content
Journal of Cardiovascular Magnetic Resonance logoLink to Journal of Cardiovascular Magnetic Resonance
. 2025 Nov 24;28(1):102015. doi: 10.1016/j.jocmr.2025.102015

A multi-dynamic low-rank deep image prior for three-dimensional real-time cardiovascular magnetic resonance imaging

Chong Chen a, Marc Vornehm b,c, Zhenyu Bu a, Preethi Chandrasekaran d, Muhammad A Sultan a, Syed M Arshad e, Yingmin Liu d, Yuchi Han f, Rizwan Ahmad a,e,
PMCID: PMC12829113  PMID: 41297766

Abstract

Purpose

To develop a reconstruction framework for three-dimensional (3D) real-time cine cardiovascular magnetic resonance (CMR) from highly undersampled data without requiring fully sampled training datasets.

Methods

We developed a multi-dynamic low-rank deep image prior (ML-DIP) framework that models spatial image content and deformation fields using separate neural networks. These sub-networks are jointly trained per scan to reconstruct the dynamic image series directly from undersampled k-space data. ML-DIP was evaluated on (i) a 3D cine digital phantom with simulated premature ventricular contractions (PVCs), (ii) 10 healthy subjects (including 2 scanned during both rest and exercise), and (iii) 12 patients with a history of PVCs. Phantom results were assessed using peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). In vivo performance was evaluated by comparing left-ventricular function quantification (against two-dimensional [2D] real-time cine) and image quality (against 2D real-time cine and binning-based five-dimensional cine [5D-Cine]).

Results

In the phantom study, ML-DIP achieved PSNR >29 dB and SSIM >0.90 for scan times as short as 2 min, while recovering cardiac motion, respiratory motion, and PVC events. In healthy subjects, ML-DIP yielded functional measurements comparable to 2D cine and higher image quality than 5D-Cine, including during exercise with high heart rates and bulk motion. In PVC patients, ML-DIP preserved beat-to-beat variations and reconstructed irregular beats, whereas 5D-Cine showed motion artifacts and information loss due to binning.

Conclusion

ML-DIP enables high-quality 3D real-time CMR with acceleration factors exceeding 1000 by learning low-rank spatial and motion representations from undersampled data, without relying on external fully sampled training datasets.

Keywords: 3D CMR, Real-time, Accelerated, Deep image prior, Arrhythmia

Graphical abstract

ga1

1. Introduction

Cardiovascular magnetic resonance imaging (CMR) is a well-established diagnostic imaging modality. CMR data are often collected slice-by-slice with electrocardiogram (ECG) gating and during breath-holds. This approach fails in patients who cannot hold their breath or have arrhythmia. For such subjects, free-breathing real-time imaging is used as a fallback option. Two-dimensional (2D) real-time imaging has progressed at a rapid pace over the past 2 decades [1], with advanced acceleration techniques enabling spatial and temporal resolutions comparable to breath-held segmented acquisitions. However, it remains limited in visualizing and modeling 3D structures due to through-plane motion, slice misregistration, and a slice thickness of 5 to 8 mm. Although 3D imaging with volumetric coverage offers advantages over 2D imaging, existing 3D imaging paradigms have significant limitations of their own. In particular, 3D methods that use prospective gating suffer from unpredictable and prolonged scan times [2], while self-gating approaches that mitigate this issue [3] remain sensitive to binning errors, which can degrade image quality [4]. For patients with arrhythmias or irregular respiratory patterns, the image quality is often compromised due to intra-bin and inter-bin motion. Even when successful, these 3D methods cannot image beat-to-beat variations, which may carry diagnostic and prognostic value [5], [6], but instead reconstruct one “typical” cardiac and/or respiratory cycle.

Extending real-time imaging to 3D would circumvent the limitations of 2D real-time imaging and 3D binning-based imaging but requires extremely high acceleration rates (R>500). One of the few efforts in this direction is the 2023 work by Sun et al., who proposed motion-resolved real-time four-dimensional (4D) flow imaging using low-rank (LR) and subspace modeling and applied it to study beat-to-beat flow variations in 10 healthy subjects and 2 patients [7]. However, this approach requires a long acquisition time of 12.18 ± 1.39 min for imaging the aorta alone. Furthermore, depending on the selected rank and matrix size, the memory requirements of this method can be prohibitive. Earlier, in 2021, Huttinga et al. introduced MR-MOTUS, which estimates respiratory and cardiac motion fields directly from undersampled data via a low-dimensional parameterization of deformation, and demonstrated its feasibility for real-time magnetic resonance (MR)-guided radiotherapy [8]. In 2022, Zou et al. presented motion-compensated smoothness regularization on manifolds (MoCo-SToRM) for lung imaging [9]. In this work, they approximated all frames in the time series as respiration-deformed versions of a single 3D template image. The deformation fields were modeled as output of a convolutional neural network (CNN) driven by low-dimensional latent vectors. Both the CNN weights and the 3D template image were jointly estimated from the undersampled k-space data. Recently, Kettelkamp et al. proposed DMoCo, which models 3D cardiac MRI volumes across motion phases as diffeomorphic deformations of a single template, and provided proof-of-concept results with radial sampling and 70 ms temporal resolution [10], [11]. However, this framework also uses a single template image and thus cannot handle contrast fluctuations, which are inevitable, e.g., due to inflow enhancement. Hamilton et al. proposed a low-rank deep image prior (LR-DIP) framework for 2D real-time imaging and then extended it to 3D real-time imaging [12] using a stack-of-spirals acquisition with anisotropic resolution. However, LR-DIP captures motion implicitly through low-rank modeling, whereas our recent 2D work shows that explicit motion modeling improves reconstruction quality and generates sharper images [13].

In this work, we propose and evaluate the multi-dynamic low-rank deep image prior (ML-DIP) framework, which models both motion and content variation using low-rank representations of the image and deformation field. Building on our prior 2D M-DIP framework [13], [14], we add low-rank modeling of deformation fields, which enables 3D imaging without requiring the memory footprint of the convolutional decoder to scale with batch size. Unlike LR-DIP [12] and the method by Sun et al. [7], ML-DIP explicitly models motion using deformation fields, making it more constrained. In contrast to the work by Kettelkamp et al. [10], [11], which directly solves for the image template and deformation basis with explicit spatial regularization, ML-DIP models both the image and deformation bases as outputs of CNNs and thus benefits from their inductive bias as an implicit regularizer. More importantly, ML-DIP models image content variation via a trainable image basis, facilitating recovery of dynamic changes beyond motion alone. These features make ML-DIP well-suited for a wide range of 3D real-time applications where both motion and contrast evolve across frames.

2. Methods

2.1. Real-time imaging in 3D

In 3D real-time imaging, the goal is to recover a series of images, each of size n1 × n2 × n3. Let x(1:T):={x(t)}t=1T represent the 3D image series, where T is the total number of frames and x(t)CN×1 is the vectorized version of the tth frame with N = n1 × n2 × n3 voxels. Throughout this paper, we will use the notation (⋅)(1:T) to represent time series of arrays or operations with T frames. For an acquisition with M measured k-space samples per frame, let y(t)CM×1, A(t)CM×N, and ϵ(t)CM×1 represent the noisy multi-coil k-space data, forward operator, and additive white Gaussian noise with variance σ2, respectively, for the tth frame. One could attempt to solve this problem using regularized least squares, i.e.,

xˆ(1:T)=argminx(1:T)t=1TA(t)x(t)y(t)22+λR(x(1:T)), (1)

where the term R(x(1:T)) represents spatial and/or temporal regularization controlled by λ≥0. However, due to the extremely high acceleration rates and large memory demands of 3D real-time imaging, directly solving the y(t)A(t)x(t)ϵ(t) problem is generally infeasible. Binning-based recovery methods address this by distributing the collected k-space data into a discrete number of cardiorespiratory bins and solving the problem in Equation (1) to reconstruct a representative cardiac and/or respiratory cycle. While these methods have been successful in many research settings [15], [16], [17], they are not real-time and cannot capture beat-to-beat variations. Additionally, these methods can degrade significantly or fail completely in the presence of frequent arrhythmias, inconsistent respiratory motion, or bulk motion.

2.1.1. Extending deep image prior to 3D imaging

Deep image prior (DIP) provides an unsupervised learning framework for solving inverse problems without requiring training data [18]. In DIP, a generative network is trained to map a random code vector to an output consistent with the measurements. A key feature of DIP is that the network structure acts as an implicit prior, eliminating the need for explicit regularization. Another key feature of DIP is that it is instance-specific, i.e., the network training is performed from scratch for each set of measurements.

A natural extension of DIP for dynamic imaging involves combining it with manifold learning. This is achieved by modeling the nonlinear mapping using a neural network:

x(t)=Gξ(z(t)), (2)

where Gξ:RK×1CN×1 is a network parameterized by ξ and z(t)RK×1 is a low-dimensional latent code vector of user-defined dimensionality K.

For given multi-coil k-space data y(t) and forward operator A(t), injecting the nonlinear mapping in Equation (2) to the data consistency term in Equation (1) leads to this optimization problem:

ξˆ,zˆ(1:T)=argminξ,z(1:T)t=1TA(t)Gξ(z(t))y(t)22. (3)

After training, an arbitrary tth frame can be recovered by xˆ(t)=Gξˆ(zˆ(t)). Recently, several studies have used DIP to learn manifolds for 2D dynamic MRI applications [19], [20]. While these approaches capture redundancy across frames through a shared network, they do not fully capture the temporal structure in a dynamic image series x(1:T). Since these approaches are not adequately constrained, our initial efforts to directly extend them to 3D real-time imaging, where the acceleration rates can be 2 orders of magnitude higher, were unsuccessful.

Recently proposed MoCo-SToRM [9] offers a more constrained approach to extending DIP to 3D real-time imaging. Instead of directly generating the more complex x(1:T), it uses the network Gξ to generate frame-specific 3-directional deformation fields ϕ(t)RN×3 and then solves the following optimization problem:

ξˆ,zˆ(1:T),xˆst=argminξ,z(1:T),xstt=1TA(t)(xstϕ(t))y(t)22,ϕ(t)=Gξ(z(t)), (4)

where xstCN×1 denotes a single static template image and “∘” denotes the spatial warping operation [21]. This framework, however, can only model motion but not other dynamics, such as contrast fluctuations. Also, this method infers xst directly and does not model it as an output of a CNN. Therefore, it does not leverage the inductive bias of CNNs but rather relies on explicit regularization of xst, which is omitted from Equation (4) for simplicity. The proposed method, described next, circumvents these limitations.

2.1.2. ML-DIP framework

ML-DIP integrates manifold learning, deformation-based motion modeling, and scalable low-rank representation into a unified framework. A high-level description of ML-DIP is provided in Fig. 1. ML-DIP generates a deformation field basis and an image basis using 2 separate CNNs. The corresponding elements from these bases are then combined to yield a frame-specific deformation field and a frame-specific composite image. The composite image is subsequently warped using the deformation field to produce an output frame. This strategy confines the dynamic component of learning to a small set of weights used to combine the basis elements. Moreover, the low-rank representation of both the deformation field and the image enables modeling of motion and image content variations (e.g., contrast), while also reducing the size of the CNN generators. Finally, since the combination weights are generated by fully connected networks with a low-dimensional input, they support learning motion and content variations through manifold learning.

Fig. 1.

Fig. 1

Overview of ML-DIP, showing the flow of information for the τth frame. A: Trainable static code vectors z˙ that serve as input to the generator Gδ. B: A series of trainable dynamic code vectors z(1:T), with a frame-specific code vector z(τ) serving as input to the generators Gω and Gν. C: Trainable static code vectors z¨ that serve as input to the generator Gβ. D: Decoder-based CNN Gδ to generate deformation field basis. E: Fully connected network Gω to generate frame-specific weights W(τ) to combine the elements of the deformation field basis. F: Fully connected network Gν to generate frame-specific weights v(τ) for combining the elements of the image basis. G: U-Net CNN Gβ to generate image basis. H: Static deformation field basis generated by Gδ. I: Linearly combining the deformation basis elements using W(τ) to generate a frame-specific deformation field. J: The 3 components of the deformation field ϕ(τ). K: Static image basis generated by Gβ. L: Linearly combining the image basis elements using v(τ) to generate a frame-specific composite image. M: The complex-valued composite image c(τ). N: Warping operation “∘” where the composite image c(τ) is deformed by ϕ(τ). O: The predicted 3D frame x˜(τ). P: The measured multi-coil k-space data y(τ). Q: The loss function that penalizes the discrepancy between A(τ)x˜(τ) and y(τ) and promotes spatial smoothness in the deformation fields. CNN convolutional neural network, ML-DIP multi-dynamic low-rank deep image prior

In ML-DIP, we jointly train 4 sub-networks and 3 sets of code vectors to generate an output frame that is consistent with the undersampled k-space data for that frame. As shown in Fig. 1, a CNN-based generator Gδ takes a static code vector z˙ as input and generates deformation field basis d(1:L1) as output, where L1 is the number of elements in the basis and d(i)RN×1 is the ith element of the deformation basis. This generator uses a decoder architecture and is parameterized by δ. We refer to this network as ConvDecoder. Another CNN-based generator Gβ takes a different static code vector z¨ as input and generates image basis b(1:L2) as output, where L2 is the number of elements in the basis and b(i)CN×1 is the ith element of the image basis. This generator uses a U-Net architecture and is parameterized by β. The outputs of Gδ and Gβ are static, i.e., a single deformation field basis and a single image basis are generated for the entire real-time image series. A small fully connected network Gω, parameterized by ω, takes frame-specific code vector z(t)RK×1 of dimensionality K as input and generates a frame-specific weight matrix W(t)RL1×3. The role of W(t) is to linearly combine elements of d(1:L1) into the frame-specific deformation field ϕ(t)=[ϕ1(t),ϕ2(t),ϕ3(t)], which is comprised of 3 components corresponding to the 3 spatial axes. The operation of W(t) on d(1:L1) can be described in terms of matrix-matrix multiplication DW(t), where the matrix DRN×L1 contains elements of d(1:L1) as its columns. Another small fully connected network Gν, parameterized by ν, takes the same frame-specific z(t)RK×1 as input and generates a frame-specific weight vector v(t)CL2×1. The role of v(t) is to linearly combine elements of b(1:L2) into the frame-specific composite image c(t). The operation of v(t) on b(1:L1) can be described in terms of matrix-vector multiplication Bv(t), where the matrix BCN×L2 contains elements of b(1:L2) as its columns. The resulting frame-specific deformation field ϕ(t) is then used to spatially warp the frame-specific composite image c(t) to generate x˜(t), which is the prediction of the tth frame.

The joint training of the 4 sub-networks and 3 code vectors is realized by solving the following optimization problem.

δˆ,βˆ,ωˆ,νˆ,z˙ˆ,z¨ˆ,zˆ(1:T)=argminδ,β,ω,ν,z˙,z¨,z(1:T)t=1TA(t)(Bv(t)c(t)DW(t)ϕ(t)x(t))y(t)22+λR(ϕ(1:T)), (5)

Here, W(t)Gω(z(t)), v(t)Gν(z(t)), d(1:L1)Gδ(z˙), and b(1:L2)Gβ(z¨) represent the outputs of 4 generators. As mentioned previously, the matrix D is generated by concatenating the elements of d(1:L1) as columns, and the matrix B is generated by concatenating the elements of b(1:L2) as columns. The term R(ϕ(1:T)) represents the spatial and/or temporal regularization applied to the deformation fields. The strength of the regularization is controlled by λ≥0. The “∘” operation represents spatial warping of the composite image. Note, the warping operation is typically realized using 3D deformation fields and 3D images and not their vectorized representations. However, for notational simplicity, we express this operation between 2 vectors, Bv(t) and DW(t).

After the network is trained, the tth 3D frame can be recovered by passing the optimized code vectors, z˙ˆ, z¨ˆ, and zˆ(t), through trained sub-networks, Gδˆ, Gβˆ, Gωˆ, and Gνˆ. This on-demand production of one or more frames obviates the need to generate or save the entire image series, which may have thousands of 3D frames for the cine acquisition performed over several minutes.

2.1.3. Implementation details of ML-DIP

As shown in Fig. 1, ML-DIP consists of 4 sub-networks and 3 sets of code vectors. The architectures of the sub-networks are reported in Appendix A, Appendix A. The input to ConvDecoder, z˙, had h = 2 channels, with each channel consisting of a real-valued 3D array. The size of the 3D array was determined by the number of upsampling steps between the input and output of ConvDecoder. Likewise, the input to U-Net, z¨, had h = 2 channels, with each channel consisting of a real-valued 3D array. The size of the 3D array was identical to the size of the target image. The entries of z˙ and z¨ were initialized independently from a uniform distribution. The sizes of the real-valued deformation basis (L1) and complex-valued image basis (L2) were set at 32 and 4, respectively. The entries of z(1:T) were initialized from the 6 principal components of the self-gating signal extracted from the repeated sampling of a central k-space line [22]. To regularize the deformation fields, R(ϕ(1:T)) in Equation (5) was chosen to be the sum of squares of the finite differences computed along the 3 spatial directions, with λ = 0.05. The sub-networks and the code vectors were jointly trained for 48000 iterations with the Adam optimizer. The batch size, which corresponds to the number of contiguous frames used in each update step, was set to 20. A cosine annealing learning rate schedule [23] was used to reduce the learning rate from its initial value of 1 × 10−3 to the final value of 2 × 10−4. The total number of learnable parameters was approximately 17 million. After training, the final network parameters and code vectors were saved. For inference, optimized z˙ˆ, z¨ˆ, and zˆ(τ1:τ2) were passed through the trained network to generate xˆ(τ1:τ2), i.e., frames in time interval τ1tτ2 for user-defined values of τ1 and τ2, such that 1≤τ1τ2T.

2.2. Experiments

To evaluate ML-DIP, we analyzed data from a 3D MRXCAT phantom [24], 10 healthy subjects at rest (2 also scanned during in-magnet exercise), and 12 patients with a history of premature ventricular contractions (PVCs). All in vivo studies were approved by the Institutional Review Board, and written informed consent was obtained. The subject characteristics of the participants are summarized in Table 1.

Table 1.

Human subjects characteristics

Healthy subjects PVC patients
Number 10 12
Age (years) 31 ± 10 48 ± 18
Sex (M/F) 4/6 4/8
BMI (kg/m2) 27.2 ± 5.3 25.1 ± 4.5
BSA (m2) 1.5 ± 0.2 1.4 ± 0.3

BMI body mass index, BSA body surface area, PVC premature ventricular contraction. Data are mean ± standard deviation or a number.

2.2.1. MRXCAT phantom

MRXCAT is a numerical simulation framework for cardiac MR imaging based on the extended cardiac-torso (XCAT) phantom [24]. It provides realistic anatomical models of the heart and thorax and supports user-defined cardiopulmonary motion patterns. Additionally, MR-specific parameters such as tissue properties, coil configurations, and noise levels can be customized in MATLAB (MathWorks, Natick, Massachusetts). To evaluate the performance of ML-DIP, we simulated a 3D cine MRXCAT phantom with an isotropic spatial resolution of 2 mm and 358 frames spanning 5 distinct respiratory cycles and 20 cardiac beats. One PVC beat was included in each respiratory cycle by shortening and altering the cardiac cycle. The 3D cine was then repeated 25 times along the temporal dimension to simulate a prolonged 5-min scan consisting of T0 = 8950 frames. To reduce computation time, the imaging volume was cropped to 110 × 112 × 92 voxels along the superior-inferior (SI), anterior-posterior, and left-right directions, respectively. Complex multicoil k-space data were generated using 8 receive coils and undersampled using ordered pseudo-radial sampling (OPRA) [25]. The SI direction was used as the frequency encoding direction (kx), anterior-posterior as the phase encoding direction (ky), and left-right as the slice encoding direction (kz). To mimic the MRI acquisition, no undersampling was applied along the kx direction. A total of 11 readouts were simulated for each 3D frame, resulting in the net acceleration rate of R = 936. Fig. 2 shows the sampling pattern used for the phantom. The sixth readout in each frame was collected along the SI orientation at kykz = 0. These central lines were subjected to bandpass filtering followed by principal component analysis (PCA) to extract 6 motion components, including 2 for respiratory and 4 for cardiac [22]. These extracted motion components were used to initialize z(1:T) in ML-DIP training.

Fig. 2.

Fig. 2

The Cartesian sampling pattern, OPRA, on a 112 × 92 grid. Here, = 1, = 2, and = 3 represent the first 3 frames, each with 11 readouts (white dots), and the “Average” represents the average of all T frames. The readout dimension kx is fully sampled and not shown. The red arrows show the acquisition order, highlighting smooth transitions from one readout to the next, even across frames. The green arrow points to the self-gating readouts at kykz = 0. ky phase encoding direction, kz slice encoding direction, OPRA ordered pseudo radial sampling

To investigate the impact of T (“scan time”) on the performance of ML-DIP, the model was trained separately for eight different values of T, i.e., T = T0 = 8950, T=45T0=7160, T=35T0=5370, T=25T0=3580, T=15T0=1790, T=110T0=895, T=125T0=358, and T190T0=100. This was achieved by truncating the original image series before training.

2.2.2. Imaging healthy subjects with ferumoxytol

ML-DIP was also validated using prospectively undersampled 3D real-time cine data collected from 10 healthy volunteers, 2 of whom were additionally scanned during in-magnet exercise at a workload of 40 W. The subjects were recruited to participate in an unrelated exercise imaging study [26], with the 3D cine acquisition added as an ancillary scan. All volunteers were imaged with an ungated spoiled gradient echo-based 3D cine sequence on a 3T scanner (MAGNETOM Vida, Siemens Healthineers, Forchheim, Germany), following ferumoxytol infusion at 4 mg/kg (Covis Pharma, Waltham, Massachusetts). The 3D cine data were acquired under free-breathing conditions for 5 min with 11 OPRA readouts per 3D frame, as shown in Fig. 2. The imaging volume was a sagittal slab covering the entire heart, with frequency encoding in the SI direction, phase encoding in the anterior-posterior direction, and slice encoding in the left-right direction. Other imaging parameters included: TE/TR of 1.2/3.1 ms; temporal resolution of 32–34 ms; acceleration rate of R = 1047; matrix size of 190 × 144 × 80; spatial resolution of 1.4–2.2 mm along frequency encoding, 1.6–2.2 mm along phase encoding, and 1.8–2.3 mm along slice encoding; and flip angle of 15. Each dataset contained approximately 9000 frames.

For comparison, each volunteer also underwent a free-breathing scan with a spoiled gradient-echo 2D real-time cine research sequence [26]. A prospectively undersampled short-axis stack covering the whole heart and one or more long-axis views was collected using variable-density golden-ratio offset sampling (R = 12) [25]. Data were reconstructed inline with Gadgetron [27] using a parameter-free compressed-sensing method [28]. The 2D scans had a spatial resolution 2.0–2.1 mm, a temporal resolution 37–39 ms, and a scan time of 6 s per slice. The 2D and 3D acquisitions were completed within 5 min of each other.

2.2.3. Imaging PVC patients post-gadolinium

To assess the ability of ML-DIP to recover irregular beats, the 3D cine acquisition was appended to the clinical exam of 12 patients with a history of frequent PVCs. Six patients exhibited irregular, frequent PVCs during the scan, 3 were predominantly in bigeminy, and 3 did not exhibit frequent PVCs during the 5-min acquisition. The patients were scanned on a 1.5T scanner (MAGNETOM Sola, Siemens Healthineers, Forchheim, Germany). Data were acquired under free-breathing conditions for 5 min using a balanced steady-state free-precession sequence with 11 OPRA readouts per 3D frame. The imaging volume was a sagittal slab covering the entire heart, with frequency encoding in the SI direction, phase encoding in the anterior-posterior direction, and slice encoding in the left-right direction. Imaging parameters included: TE/TR = 1.3/3.1 ms, temporal resolution = 32–34 ms, acceleration rate = R = 1047, matrix size = 190 × 144 × 80, spatial resolution = 1.4–1.8 mm (frequency), 1.5–1.8 mm (phase), and 2.0 mm (slice), and flip angle = 34–40. Each dataset contained approximately 9000 frames. For comparison, 2D real-time imaging was also performed using a balanced steady-state free-precession sequence with 3 s per slice. The spatial and temporal resolutions of the 2D scans were similar to those used in the healthy subject study. Unlike the previous study, ferumoxytol was not used in this patient study, and 3D acquisition occurred approximately 20–30 min after administration of a gadolinium-based contrast agent (Gadovist, Bayer AG, Berlin, Germany).

2.3. Data processing and analysis

2.3.1. Preprocessing

To reduce computation time for in vivo studies, the k-space data in the kx-ky-kz domain were first transformed to the spatial domain along the readout (kx) dimension using a 1D fast Fourier transform (FFT). The data were then cropped along this dimension and transformed back to the kx-ky-kz domain using an inverse 1D FFT. This step was performed by presenting users with a time-averaged sagittal image and prompting them to select 2 points, one above and one below the heart. Although not required, this subject-specific cropping reduced computation time by limiting the readout dimension. To further accelerate computation, the physical coils were compressed to fifteen virtual coils using PCA [29] and then to eight using region-optimized virtual coils [30]. Since region-optimized coil compression has the tendency to select virtual coils that are mostly noise but have a favorable signal-to-interference ratio, PCA-based compression was performed first as a denoising step. Coil sensitivity maps of the eight virtual coils were then estimated from the time-averaged k-space using ESPIRiT [31].

2.3.2. Image reconstruction

For each dataset, training was performed on a single RTX 6000 Ada (Nvidia, Santa Clara, California). Depending on the size of the imaging matrix, training time was approximately 7 to 10 h. For the MRXCAT phantom, 100 consecutive frames were generated post-training. The inference time to generate these frames was approximately 2–3 s. For the data from healthy subjects and patients, 200 consecutive frames were generated post-training, with an inference time of 5–8 s. For comparison, the same 3D cine data were binned into 20 cardiac and 4 respiratory phases and reconstructed using compressed sensing [3], with regularization applied along the spatial, respiratory, and cardiac dimensions. We refer to this method as 5D-Cine. The reconstruction time for 5D-Cine was approximately 2 hours on the same workstation. To improve the feasibility of 5D-Cine in PVC patients, arrhythmia rejection was applied by discarding k-space data from beats that deviated significantly from the average beat length [32]. Only the expiratory phase was analyzed from 5D-Cine, as it provided the highest image quality among the 4 respiratory phases.

2.3.3. Image analysis

For the phantom data, where the noiseless ground truth was available, image quality of the generated frames was assessed using peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) for all eight values of T. For the data collected from human subjects, both ML-DIP and 5D-Cine results were interpolated into standard short-axis and long-axis views using slice position information from the 2D real-time cine data. To evaluate the image quality of 5D-Cine, 2D real-time cine, and ML-DIP, one short-axis and one long-axis view from each subject were blindly scored by 2 experienced Level-3 CMR-trained cardiologists. The 3 cine series were presented as movies on a single slide, with the order of 5D-Cine, 2D real-time cine, and ML-DIP randomized. Each reader assigned a score from 1 (worst) to 5 (best) to each cine based on overall image quality: 1 - Non-diagnostic, 2 - Poor, 3 - Fair, 4 - Good, 5 - Excellent. The interpolated short-axis stacks from 5D-Cine and ML-DIP were converted to digital imaging and communications in medicine (DICOM) format and analyzed in SuiteHEART (NeoSoft, Pewaukee, Wisconsin) along with the 2D real-time stack. Left-ventricular end-diastolic volume (EDV), end-systolic volume (ESV), stroke volume (SV), and ejection fraction (EF) were computed from 2D real-time cine, 5D-Cine, and ML-DIP using the short-axis stacks. Quantification was performed only for subjects whose average image quality score across both readers and both views was at least 3.

3. Results

Table 2 shows image quality metrics from the MRXCAT phantom for different values of T. As expected, image quality decreased with shorter scan durations due to reduced k-space coverage. However, there was no significant drop in PSNR or SSIM when T was reduced from 8950 (5 min) to 3580 (2 min), and only a modest drop in the metrics at T = 1790 (1 min). A more substantial decline in image quality metrics was observed when T was at or below 895 (30 s). Fig. 3 shows representative images at different values of T. These images show that ML-DIP effectively recovered both respiratory and cardiac motion, including PVC beats. A PVC beat, marked by yellow arrows, is evident in the space-time (x-t) profiles. Only the images at T = 358 and T = 100 showed noticeable artifacts and noise amplification. These results suggest that ML-DIP can be effective at shorter acquisition times. A movie, Video S1, corresponding to Fig. 3, is provided in the Supplementary material.

Table 2.

PSNR and SSIM values for eight different numbers of total frames, T

T = 8950 T = 7160 T = 5370 T = 3580 T = 1790 T = 895 T = 358 T = 100
PSNR (dB) 29.9 30.1 29.9 29.5 29.2 27.9 24.7 19.9
SSIM 0.95 0.95 0.94 0.92 0.89 0.82 0.71 0.54

PSNR peak signal-to-noise ratio, SSIM structural similarity index measure, T number of frames. Data are values of image quality metrics. These values of T, from left to right, correspond to the scan times of 300, 240, 180, 120, 60, 30, 12, and 3.4 s, respectively

Fig. 3.

Fig. 3

Representative images from MRXCAT reconstructed using ML-DIP. Time profiles along the dashed green lines are shown to the right of the images. Each x-t profile spans 100 frames. A simulated PVC beat is highlighted by the yellow arrow. Here, T represents the number of frames and the numbers in the parentheses represent scan times, assuming a repetition time is similar to the ones reported for human studies. GT ground truth, ML-DIP multi-dynamic low-rank deep image prior, MRXCAT magnetic resonance extended cardiac-torso, PVC premature ventricular contraction

Appendix A, Appendix A in the Supplementary material highlights the advantage of using a frame-specific composite image in ML-DIP. Relying on a single fixed template resulted in visible image distortions, as highlighted by the yellow arrows. Because the fixed-template approach performed poorly compared to ML-DIP, it was excluded from further comparisons. Fig. 4 shows representative ML-DIP reconstructions from eight healthy volunteers scanned at rest. The corresponding x-t profiles highlight that ML-DIP preserves both cardiac and respiratory motions, closely following the self-gating signals. Fig. 5 shows results from 2 additional volunteers scanned both at rest and during in-magnet exercise. During exercise, faster heart rates and exaggerated breathing patterns were captured without visible degradation in image quality. Fig. 6 shows representative results from 6 of the 12 PVC patients. Despite lower blood-myocardium contrast, ML-DIP successfully captured beat-to-beat variations, including the timing and morphology of PVCs. PVC beats were easily identified on the reconstructed x-t profiles and corroborated by the self-gating signals.

Fig. 4.

Fig. 4

Representative ML-DIP images in sagittal orientation from 8 subjects imaged at rest. Time profiles along the dashed green lines are shown to the right of the images. Each x-t profile spans 200 frames (6.6 s). The red and cyan curves represent self-gating-based respiratory and cardiac signals, respectively. ML-DIP multi-dynamic low-rank deep image prior

Fig. 5.

Fig. 5

Representative ML-DIP images in sagittal orientation from 2 subjects imaged at rest and during in-magnet exercise. Time profiles along the dashed green lines are shown to the right of the images. Each x-t profile spans 200 frames (6.6 s). The red and cyan curves represent self-gating-based respiratory and cardiac signals, respectively. Significantly faster heart rates are observed during exercise. ML-DIP multi-dynamic low-rank deep image prior

Fig. 6.

Fig. 6

Representative ML-DIP images in sagittal orientation from 6 out of 12 PVC patients. Space-time (x-t) profiles along the dashed green lines are shown to the right of the images. Each x-t profile spans 200 frames (6.6 s). The red and cyan curves represent self-gating-based respiratory and cardiac signals, respectively. PVC beats are indicated by yellow arrows. The compensatory pause is also visible after some of the PVC beats. One of the patients (bottom-left) was in bigeminy, evident in the x-t profile, and another patient (bottom-right) did not experience PVCs during the scan. ML-DIP multi-dynamic low-rank deep image prior, PVC premature ventricular contraction

In healthy subjects, blind scoring by cardiologists consistently yielded higher scores for ML-DIP, as summarized in Table 3. The advantage of ML-DIP over 5D-Cine was more pronounced during exercise, where 5D-Cine exhibited visible motion artifacts due to unaccounted torso movement. In PVC patients, where ferumoxytol was not used, 2D real-time scored higher than ML-DIP, which can be attributed to lower contrast in 3D imaging from blood pool saturation [33]. Nonetheless, ML-DIP received an average score of 4.50, with the lowest score of 3.75 for PVC #3. Despite arrhythmia rejection, 5D-Cine performed poorly in PVC patients, with 8 subjects scoring below 3.00 and 4 scoring below 2.00. Only 4 PVC patients—3 without arrhythmia during acquisition and 1 in stable bigeminy—received a score of 3.00 or higher. Fig. 7 shows some of the short-axis images scored by expert readers. In healthy subjects, ML-DIP produced sharp, motion-artifact-free images, even during exercise. In PVC patients, ML-DIP reconstructed PVC beats, including the compensatory pauses following PVCs in one subject and alternating premature beats consistent with bigeminy in another. Beat-to-beat variations observed with ML-DIP were also consistent with the ECG traces, which are superimposed on the temporal profiles of the 2 PVC subjects in Fig. 7. In contrast, 5D-Cine not only removed beat-to-beat variations but also exhibited extensive motion artifacts and blurring due to inconsistent cardiac binning. A movie, Video S2, corresponding to Fig. 7, is shown in the Supplementary material. To highlight the improvement in blood-myocardium contrast from ferumoxytol, Video S3 in the Supplementary material presents 2 ML-DIP cine series: one from a healthy subject scanned with ferumoxytol at 3T and the other from a PVC patient scanned post-gadolinium at 1.5T.

Table 3.

Image quality scoring for the human subject studies

Dataset 5D-Cine 2D real-time ML-DIP
Vol. #1 4.25 4.25 4.75
Vol. #2 4.00 4.00 5.00
Vol. #3 5.00 4.25 5.00
Vol. #4 4.75 4.00 4.50
Vol. #5 4.00 4.25 5.00
Vol. #6 5.00 4.50 5.00
Vol. #7 3.50 4.00 5.00
Vol. #8 4.75 4.25 4.50
Vol. #9 5.00 4.50 5.00
Vol. #10 4.50 4.25 5.00
Vol. Avg. 4.48 4.23 4.88
Vol. #9-Ex. 3.00 4.00 4.25
Vol. #10-Ex. 3.25 4.50 5.00
Vol. Ex. Avg. 3.13 4.25 4.63
PVC #1 1.00 4.00 4.00
PVC #2 3.75 4.75 4.00
PVC #3 2.25 4.50 3.75
PVC #4 2.00 4.25 4.25
PVC #5 3.00 5.00 4.50
PVC #6 2.75 5.00 5.00
PVC #7 1.75 4.75 4.75
PVC #8 4.00 5.00 5.00
PVC #9 1.25 5.00 4.50
PVC #10 3.00 4.00 4.50
PVC #11 2.75 4.75 4.50
PVC #12 1.00 5.00 4.75
PVC Avg. 2.38 4.67 4.50
Overall Avg. 3.31 4.45 4.67

Avg average, Ex exercise, PVC premature ventricular contraction, Vol volunteer. Data are mean values of image quality scores. For an individual dataset, each number represents an average over 2 cardiac views and 2 expert readers. Here, “Vol. Avg.” represents an average over 10 healthy volunteers imaged at rest, “Vol. Ex. Avg.” represents an average over 2 healthy volunteers imaged during exercise, “PVC Avg.” represents an average over 12 PVC patients, and “Overall Avg.” represents an average over all 24 cases. In each row, the highest number is highlighted in bold font. PVC #5, #7, and #11 were predominantly in bigeminy, and #2, #8, and #10 did not exhibit frequent arrhythmias during the acquisition

Fig. 7.

Fig. 7

Representative short-axis images from 5D-Cine, 2D real-time, and ML-DIP for a healthy subject at rest, a healthy subject during exercise, and 2 patients. Space-time (x-t) profiles along the dashed green lines are shown to the right of the images. The ML-DIP x-t profiles span 200 frames (6.6 s), while the x-t profiles span 3 to 6 s for 2D real-time cine and one cardiac cycle for 5D-Cine. The 3D reconstructions from 5D-Cine and ML-DIP were interpolated along the 2D plane defined by the 2D real-time acquisition. Yellow traces show the ECG signal that was synchronously collected with the 3D acquisition. R-waves from sinus and PVC beats are marked with ‘*’ and ‘**,’ respectively. The yellow arrows indicate PVCs seen on the cine images. Artifacts due to uncompensated motion are observed in 5D-Cine, especially in the second, third, and fourth rows. Because 2D and 3D scans were performed separately, minor shifts in subject position are seen in some cases. 5D-Cine cardiorespiratory-binning-based volumetric cine, ECG electrocardiogram, ML-DIP multi-dynamic low-rank deep image prior, PVC premature ventricular contraction

The cardiac function quantification results are summarized in Fig. 8. In healthy subjects, LV functional measurements from both 5D-Cine and ML-DIP showed excellent agreement with those from 2D real-time cine, including during exercise. In PVC patients, functional analysis was feasible for all 12 subjects using ML-DIP, but only for 4 subjects using 5D-Cine due to poor image quality. Across all 24 acquisitions (healthy + patients), the mean absolute difference between ML-DIP and 2D real-time was 8.7 mL for EDV, 6.2 mL for ESV, 5.8 mL for SV, and 2.6% for EF. Pearson correlation coefficients between ML-DIP and 2D real-time were >0.9 for all metrics. All volumetric measurements were limited to the first dominant sinus contraction, identified as a large drop in LV volume. In PVC patients, beat-to-beat or arrhythmic-beat quantification was not feasible with 2D real-time imaging because that would require substantially longer scan times per slice and manual alignment of matching beats across slices [34]. In contrast, ML-DIP yields temporally coherent 3D volumes, enabling visualization of beat-to-beat variations in cardiac function. Fig. 9 shows representative examples demonstrating these variations in 2 PVC patients together with the corresponding ECG traces.

Fig. 8.

Fig. 8

LV quantification from 5D-Cine and ML-DIP, with 2D real-time used as the reference. In the case of 5D-Cine, only 4 (out of 12) PVC patients are included because the image quality from 8 other patients was not adequate for analysis. (a) Correlation plot. (b) Bland-Altman plot, where “difference” represents volumetric reconstruction - 2D. 5D-Cine cardiorespiratory-binning-based volumetric cine, EDV end-diastolic volume, EF ejection fraction, ESV end-systolic volume, LoA limits of agreement, LV left ventricle, MAD mean absolute difference, ML-DIP multi-dynamic low-rank deep image prior, PVC premature ventricular contraction, SV stroke volume

Fig. 9.

Fig. 9

LV volumes (dashed lines) across ML-DIP frames from 2 different PVC patients in (a) and (b). The corresponding ECG traces (solid lines), expressed in arbitrary units, are also shown at the bottom. R-waves corresponding to sinus beats are marked with ‘*,’ and those corresponding to PVCs are marked with ‘**.’ ECG electrocardiogram, LV left ventricle, ML-DIP multi-dynamic low-rank deep image prior, PVC premature ventricular contraction

4. Discussion

Three-dimensional real-time imaging has the potential to offer a paradigm shift for CMR imaging. However, due to very high acceleration rates (R>1000), 3D real-time imaging has not been previously possible. In this work, we develop a scan-specific learning-based framework, called ML-DIP, that can facilitate 3D real-time imaging from a free-breathing ungated scan. A key innovative feature of ML-DIP is its ability to model multiple, disparate image dynamics using scalable, low-rank representations.

Our phantom results demonstrate that ML-DIP can recover image series with high values of PSNR and SSIM. As evident from Table 2, image quality degrades as the value of T gets smaller, with the images from T = 358 and T = 100 datasets showing significant noise amplification. This is expected because the learning in ML-DIP is based on the entire time series and not the individual frames. Since each frame contributes a unique, complementary sampling pattern, smaller values of T leave most indices in k-space unsampled for any one motion state. Nonetheless, the PSNR and SSIM values of ML-DIP stay within 0.5 dB and 0.03, respectively, of the best values until T = 3580, which corresponds to a 2-min scan. Although we have performed all subsequent in vivo studies with a fixed acquisition time of 5 min, this phantom study suggests that there is a margin to reduce the acquisition time to well below 5 min.

Our healthy subject study demonstrates that ML-DIP can generate high-quality images while preserving cycle-to-cycle variations in respiratory and cardiac motion. For images collected at rest, all methods received a high image quality score of 4.00 or higher. However, images from ML-DIP were visibly sharper than those from 5D-Cine and even 2D cine, where the former suffers from blurring due to intra-bin motion and the latter from inhomogeneous blood pool intensity due to the mixing of excited and unexcited blood in the presence of ferumoxytol. The quality gap between ML-DIP and 5D-Cine was more pronounced during exercise. In this setting, ML-DIP captured cardiac, respiratory, and bulk motion due to pedaling on an ergometer in real time, whereas 5D-Cine reconstructed only a representative cardiorespiratory cycle, with unaccounted bulk motion manifesting as artifacts. Nonetheless, due to high blood-myocardium contrast, 5D-Cine received scores of at least 3.00 in all cases.

The 3D imaging of PVC subjects is more challenging due to lower blood-myocardium contrast without ferumoxytol, reduced signal-to-noise ratio at lower field strength, and irregular heart rhythms. Despite these challenges, ML-DIP successfully captured both respiratory and cardiac dynamics, including PVC beats. Although the self-gating signal is less precise in subjects with irregular breathing or arrhythmias, it remains sensitive to detecting irregular events. We therefore used self-gating as well as synchronously acquired ECG as surrogates to validate the timing and consistency of the reconstructed PVC dynamics. We observed that the PVC occurrences in the ML-DIP reconstructions aligned temporally with both the self-gating cardiac signal (Fig. 6) and the ECG signal (Fig. 7), and their temporal patterns matched those observed in 2D real-time images. However, 2D imaging captures arrhythmias in only a limited number of slices and may miss events occurring outside those planes or acquisition windows. ML-DIP overcomes this limitation by providing volumetric coverage with consistent temporal sampling, enabling whole-heart assessment of arrhythmias from a single scan. Furthermore, because ML-DIP does not rely on binning or arrhythmia rejection, it preserves true beat-to-beat variation and avoids the artifacts and information loss associated with retrospective gating. While larger validation studies are needed, this preliminary evidence suggests that ML-DIP can reliably recover complex, subject-specific motion patterns in challenging patient populations.

5. Limitations

The current study has several limitations. First, our sample size is small and includes only one type of arrhythmia, i.e., PVC. Second, completely non-contrast acquisitions were not considered and are expected to offer lower blood-myocardium contrast. Third, the reconstruction time for ML-DIP is long, which limits translational potential in the short term. However, one might be able to accelerate ML-DIP by pre-training the ConvDecoder and/or U-Net on the coarse results obtained from 5D-Cine before jointly training all 4 sub-networks in an end-to-end fashion. Another avenue to accelerate ML-DIP is to adopt a hierarchical approach where the spatial and temporal resolutions of the images are progressively increased over iterations. Fourth, additional validation is needed to confirm that ML-DIP preserves strain patterns and regional wall motion abnormalities, which are important for detecting subtle functional impairments. Fifth, all the presented studies were conducted at 32–34 ms temporal resolution; the impact of changing temporal resolution on the quality of ML-DIP reconstructions needs to be further explored. Sixth, the current architectures used in ML-DIP and the values of hyperparameters are not fully optimized. This is primarily due to the long reconstruction times. Future studies of ML-DIP will include a larger sample size and a more diverse set of CMR applications.

6. Conclusions

We have presented and evaluated a scan-specific framework, called ML-DIP, for 3D real-time CMR. A key feature of ML-DIP is its ability to model both motion and contrast changes. The low-rank modeling for both the image and deformation representations makes the network architecture simpler and easier to train. Phantom and in vivo results demonstrate the potential of ML-DIP to preserve real-time dynamics from highly undersampled 3D data.

Author contributions

C. Chen implemented ML-DIP and generated initial results. M. Vornehm assisted with implementation and manuscript review. Z. Bu reconstructed the additional patient data and generated updated figures. M. Sultan provided feedback on improving and optimizing ML-DIP and reviewed the manuscript. M. Arshad assisted with data acquisition and curation. Y. Liu performed pulse sequence programming and scanned the subjects. P. Chandrasekaran assisted with volunteer recruitment, data acquisition, and image analysis. Y. Han contributed to experiment design and result interpretation. R. Ahmad supervised all aspects of the study.

Ethics approval and consent

For the human subject data, approval was granted by the Institutional Review Board (IRB) at The Ohio State University (2020H0402 and 2019H0076). Informed consent to participate in the study and publish results was obtained from all individual participants.

Declaration of competing interests

The authors declare no competing interests.

Availability of data and materials

ML-DIP code and sample dataset are available at https://github.com/OSU-CMR/ml-dip.

Acknowledgements

This work was funded by NIH grants R01-EB029957, R01-HL151697, and R01-HL148103.

Footnotes

Appendix A

Supplementary data associated with this article can be found in the online version at doi:10.1016/j.jocmr.2025.102015.

Appendix A. Supplementary material

Supplementary material

mmc1.pptx (72MB, pptx)

.

References

  • 1.Contijoch F., Rasche V., Seiberlich N., Peters D.C. The future of CMR: all-in-one vs. real-time CMR (Part 2) J Cardiovasc Magn Reson. 2024;26(1) doi: 10.1016/j.jocmr.2024.100998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Goo H.W. Comparison between three-dimensional navigator-gated whole-heart MRI and two-dimensional cine MRI in quantifying ventricular volumes. Korean J Radiol. 2018;19(4):704–714. doi: 10.3348/kjr.2018.19.4.704. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Feng L., Coppo S., Piccini D., Yerly J., Lim R.P., Masci P.G., et al. 5D whole-heart sparse MRI. Magn Reson Med. 2018;79(2):826–838. doi: 10.1002/mrm.26745. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Arshad S.M., Potter L.C., Chen C., Liu Y., Chandrasekaran P., Crabtree C., et al. Motion-robust free-running volumetric cardiovascular MRI. Magn Reson Med. 2024;92(3):1248–1262. doi: 10.1002/mrm.30123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.DiCarlo A.L., Haji-Valizadeh H., Passman R., Greenland P., McCarthy P., Lee D.C., et al. Assessment of beat-to-beat variability in left atrial hemodynamics using real time phase contrast MRI in patients with atrial fibrillation. J Magn Reson Imaging. 2023;58(3):763–771. doi: 10.1002/jmri.28550. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Alhede C., Higuchi S., Hadjis A., Bibby D., Abraham T., Schiller N.B., et al. Premature ventricular contractions are presaged by a mechanically abnormal sinus beat. Clin Electrophysiol. 2022;8(8):943–953. doi: 10.1016/j.jacep.2022.05.005. [DOI] [PubMed] [Google Scholar]
  • 7.Sun A., Zhao B., Zheng Y., Long Y., Wu P., Wang B., et al. Motion-resolved real-time 4D flow MRI with low-rank and subspace modeling. Magn Reson Med. 2023;89(5):1839–1852. doi: 10.1002/mrm.29557. [DOI] [PubMed] [Google Scholar]
  • 8.Huttinga N.R., Bruijnen T., van den Berg C.A., Sbrizzi A. Nonrigid 3D motion estimation at high temporal resolution from prospectively undersampled k-space data using low-rank MR-MOTUS. Magn Reson Med. 2021;85(4):2309–2326. doi: 10.1002/mrm.28562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zou Q., Torres L.A., Fain S.B., Higano N.S., Bates A.J., Jacob M. Dynamic imaging using motion-compensated smoothness regularization on manifolds (MoCo-SToRM) Phys Med Biol. 2022;67(14) doi: 10.1088/1361-6560/ac79fc. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kettelkamp J., Romanin L., Piccini D., Priya S., Jacob M. International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer,; 2023. Motion compensated unsupervised deep learning for 5D MRI; pp. 419–427. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Kettelkamp J.W., Romanin L., Priya S., Jacob M. Diffeomorphic motion-compensated (DMoCo) cardiac MRI. J Cardiovasc Magn Reson. 2025;27 [Google Scholar]
  • 12.Hamilton J., Da Cruz G.L., Seiberlich N. 3D free-breathing ungated cine imaging at 1.5T and 0.55T using a time- and partition-dependent deep image prior. J Cardiovasc Magn Reson. 2024;26 [Google Scholar]
  • 13.Vornehm M., Chen C., Sultan M.A., Arshad S.M., Han Y., Knoll F., et al. Multi-dynamic deep image prior for cardiac MRI. Magn Reson Med. 2025;94(6):2668–2679. doi: 10.1002/mrm.70000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.M. Vornehm, C. Chen, M.A. Sultan, S.M. Arshad, F. Knoll, and R. Ahmad, Motion-guided deep image prior for dynamic cardiac MRI, In: Proceedings of the International Society for Magnetic Resonance in Medicine (ISMRM)), Honolulu, HI, USA, 2025. [DOI] [PMC free article] [PubMed]
  • 15.Ma L.E., Yerly J., Piccini D., DiSopra L., Roy C.W., Carr J.C., et al. 5D flow MRI: a fully self-gated, free-running framework for cardiac and respiratory motion-resolved 3D hemodynamics. Radio Cardiothorac Imaging. 2020;2(6) doi: 10.1148/ryct.2020200219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Di Sopra L., Piccini D., Coppo S., Stuber M., Yerly J. An automated approach to fully self-gated free-running cardiac and respiratory motion-resolved 5D whole-heart MRI. Magn Reson Med. 2019;82(6):2118–2132. doi: 10.1002/mrm.27898. [DOI] [PubMed] [Google Scholar]
  • 17.Roy C.W., Di Sopra L., Whitehead K.K., Piccini D., Yerly M., Heerfordt J., et al. Free-running cardiac and respiratory motion-resolved 5D whole-heart coronary cardiovascular magnetic resonance angiography in pediatric cardiac patients using ferumoxytol. J Cardiovasc Magn Reson. 2022;24(1):39. doi: 10.1186/s12968-022-00871-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.D. Ulyanov, A. Vedaldi, and V. Lempitsky, Deep image prior, In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, 9446–9454.
  • 19.Yoo J., Jin K.H., Gupta H., Yerly J., Stuber M., Unser M. Time-dependent deep image prior for dynamic MRI. IEEE Trans Med Imaging. 2021;40(12):3337–3348. doi: 10.1109/TMI.2021.3084288. [DOI] [PubMed] [Google Scholar]
  • 20.Zou Q., Ahmed A.H., Nagpal P., Kruger S., Jacob M. Dynamic imaging using a deep generative SToRM (Gen-SToRM) model. IEEE Trans Med Imaging. 2021;40(11):3102–3112. doi: 10.1109/TMI.2021.3065948. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, Spatial transformer networks, In: Advances in Neural Information Processing Systems, 28, 2015.
  • 22.Chen C., Liu Y., Simonetti O.P., Tong M., Jin N., Bacher M., et al. Cardiac and respiratory motion extraction for MRI using Pilot Tone–a patient study. Int J Cardiovasc Imaging. 2024;40(1):93–105. doi: 10.1007/s10554-023-02966-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.I. Loshchilov and F. Hutter, SGDR: Stochastic gradient descent with warm restarts, 2017.
  • 24.Wissmann L., Santelli C., Segars W.P., Kozerke S. MRXCAT: realistic numerical phantoms for cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2014;16(1):63. doi: 10.1186/s12968-014-0063-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.M. Joshi, A. Pruitt, C. Chen, Y. Liu, and R. Ahmad, Technical report (v1.0)–Pseudo-random Cartesian sampling for dynamic MRI, 2022.
  • 26.Chandrasekaran P.S., Chen C., Liu Y., Arshad S.M., Crabtree C., Tong M., et al. Accelerated real-time cine and flow under in-magnet staged exercise. J Cardiovasc Magn Reson. 2025;27(1) doi: 10.1016/j.jocmr.2025.101894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hansen M.S., Sørensen T.S. Gadgetron: an open source framework for medical image reconstruction. Magn Reson Med. 2013;69(6):1768–1776. doi: 10.1002/mrm.24389. [DOI] [PubMed] [Google Scholar]
  • 28.Chen C., Liu Y., Schniter P., Jin N., Craft J., Simonetti O., et al. Sparsity adaptive reconstruction for highly accelerated cardiac MRI. Magn Reson Med. 2019;81(6):3875–3887. doi: 10.1002/mrm.27671. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Buehrer M., Pruessmann K.P., Boesiger P., Kozerke S. Array compression for MRI with large coil arrays. Magn Reson Med. 2007;57(6):1131–1139. doi: 10.1002/mrm.21237. [DOI] [PubMed] [Google Scholar]
  • 30.Kim D., Cauley S.F., Nayak K.S., Leahy R.M., Haldar J.P. Region-optimized virtual ROVir coils: localization and/or suppression of spatial regions using sensor-domain beamforming. Magn Reson Med. 2021;86(1):197–212. doi: 10.1002/mrm.28706. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Uecker M., Lai P., Murphy M.J., Virtue P., Elad M., Pauly J.M., et al. ESPIRiT—an eigenvalue approach to autocalibrating parallel MRI: where SENSE meets GRAPPA. Magn Reson Med. 2014;71(3):990–1001. doi: 10.1002/mrm.24751. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Pruitt A., Rich A., Liu Y., Jin N., Potter L., Tong M., et al. Fully self-gated whole-heart 4D flow imaging from a 5-minute scan. Magn Reson Med. 2021;85(3):1222–1236. doi: 10.1002/mrm.28491. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Nezafat R., Herzka D., Stehning C., Peters D.C., Nehrke K., Manning W.J. Inflow quantification in three-dimensional cardiovascular MR imaging. J Magn Reson Imaging. 2008;28(5):1273–1279. doi: 10.1002/jmri.21493. [DOI] [PubMed] [Google Scholar]
  • 34.Contijoch F., Rogers K., Rears H., Shahid M., Kellman P., Gorman J., III, et al. Quantification of left ventricular function with premature ventricular complexes reveals variable hemodynamics. Circ Arrhythmia Electrophysiol. 2016;9(4) doi: 10.1161/CIRCEP.115.003520. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material

mmc1.pptx (72MB, pptx)

Data Availability Statement

ML-DIP code and sample dataset are available at https://github.com/OSU-CMR/ml-dip.


Articles from Journal of Cardiovascular Magnetic Resonance are provided here courtesy of Elsevier

RESOURCES