Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 3.
Published in final edited form as: Ultrasonics. 2026 Mar 27;165:108081. doi: 10.1016/j.ultras.2026.108081

Deep Learning for scaling large-aperture photoacoustic computed tomography : From single fingers to the human hand

Seongwook Choi 1, Katherine W Ferrara 1,*
PMCID: PMC13228181  NIHMSID: NIHMS2164273  PMID: 41990474

Abstract

Photoacoustic Computed Tomography (PACT) leverages the photoacoustic effect for high-resolution anatomical and molecular imaging. We developed an advanced PACT system using eight conventional linear arrays arranged in a half-ring geometry, achieving a balance between cost-efficiency and enhanced image quality through a large-aperture detection setup. Although this large-aperture PACT system provides high-quality imaging for in-vivo human applications, it is susceptible to optical shadowing and misalignment issues between optical paths and detection planes, particularly during complex and large-target imaging, such as imaging of the human hand. These issues can lead to degraded image quality. To address these issues, we implemented an encoder-decoder structure-based deep learning (DL) enhancement strategy. The DL model was initially trained using a paired PACT single-finger dataset, which included images obtained with full detection using all eight transducers and those with low detection using fewer transducers or elements. For human-hand PACT imaging, the DL-enhanced system, trained exclusively with the single-finger dataset, effectively mitigated the image quality issues by improving contrast-to-noise ratios and the clarity of vessel structures. These findings validate the efficacy of the DL-enhanced PACT system for complex anatomical imaging applications, such as diagnosing peripheral arterial disease.

Keywords: Photoacoustic computed tomography, Deep learning

1. Introduction

Photoacoustic computed tomography (PACT) has emerged as a powerful non-invasive imaging modality, exploiting the photoacoustic (PA) effect to achieve high-resolution visualization of anatomical structures with excellent contrast [1–8]. In recent years, significant efforts have been dedicated to optimizing PACT systems through innovative transducer array configurations, notably arc, ring, and hemispherical shapes, each designed to enhance image quality taking advantage of large-aperture detection and broaden the field of view (FOV) [9,10]. For optimal performance of these PACT systems, the laser illumination configuration must be meticulously considered. Common geometries for laser illumination include: (1) positioning optical fibers adjacent to the transducers, ensuring the transducers directly face the illuminated region [11,12]; (2) illuminating the target from the opposite side relative to the transducers [13]; and (3) configuring the optical illumination at a 90-degree angle to the transducers [14], which also could utilize a diffuser to uniformly disperse the laser light across the imaging target [15,16]. Large targets, such as the breast and human limbs (i.e., arms and legs), are well-suited for large-aperture PACT systems, which are more capable of capturing comprehensive images of the whole region compared to small-aperture PACT systems [17–19].

Peripheral arterial disease (PAD) is a prevalent circulatory disorder characterized by arterial narrowing that limits blood flow to the extremities, underscoring the need for accurate and non-invasive vascular imaging for diagnosis and treatment planning [20]. For PAD, imaging the vascular structures of the fingertips and toes is crucial for monitoring the progression or prognosis of diseases such as diabetes. Therefore, for the diagnosis and monitoring of PAD, PACT is a promising tool that leverages non-invasive high-resolution imaging capabilities without the need for contrast agents [1,20,21]. For PACT systems, optical illumination uniformity strongly depends on the geometric relationship between the target surface and the illumination configuration. Targets with relatively smooth and convex surfaces, such as single limbs, tend to receive more homogeneous light delivery when positioned near the optical focus. However, when the target geometry becomes more complex or spatially extended, such as in multi-finger hand imaging, parts of the surface inevitably fall outside the optimal illumination region, resulting in non-uniform fluence distribution and optical shadowing.

To address these challenges, we propose a deep-learning enhanced PACT system based on eight conventional linear arrays. First, the developed PACT system was configured with eight transducer arrays in a half-ring geometry. In addition to geometric considerations, the cost and accessibility of transducer hardware play a critical role in clinical translation. Conventional commercial array transducers are substantially more affordable and widely available than customized ring, half-ring, or arc arrays, motivating increasing interest in achieving high-performance PACT using standard clinical ultrasound hardware [22–28]. The chosen configuration remarkably minimizes the limited-view effect commonly encountered in PACT, ensuring comprehensive image acquisition and facilitating extensive FOV coverage suitable for applications such as human-limb imaging [29]. To further address the optical shadowing problem and improve imaging quality for more complex structures, we integrated a deep-learning (DL) enhancement approach. Nowadays, like other biomedical imaging modalities, the field of PACT has widely adopted DL techniques for image enhancement, segmentation, and classification [11,30–33]. Specifically, we trained the neural network using paired datasets in which high-quality PA images, reconstructed from the full set of 1,024 detection elements (128 elements × 8 arrays), served as the reference.

The corresponding input images were generated under two complementary degradation schemes. First, we created 4-view datasets by selecting only four of the eight transducer arrays. These four views were either chosen as a cluster (four consecutive arrays) to emulate limited angular coverage, or randomly selected across the eight arrays to represent irregular detector availability. Second, we incorporated element-level sparse datasets in which the reconstruction used only a fraction of the total 1,024 elements. To systematically model different detector sparsity levels, we generated inputs using 1/8, 1/16, and 1/32 of the elements (i.e., 128, 64, and 32 active elements), each producing progressively degraded PA images. The sampled inputs are used to approximate image degradation patterns that arise from insufficient detection coverage or illumination–detection mismatch when imaging larger and more complex targets, such as the human hand. By leveraging DL, the network improves the image reconstruction quality, compensating for the shortcomings related to the physical configuration and illumination geometry. Once trained, the model was first validated on single-finger images, demonstrating its capability to enhance image quality to match the eight-detector standard. Subsequently, we applied the trained neural network to larger complex targets, specifically human hand imaging. This step aimed to test the model’s robustness and effectiveness in enhancing the image quality of complex structures with varying illumination and detection challenges. The results validated the DL-enhanced PACT system’s ability to mitigate optical shadowing and improve structural information capture, thus extending the application range of our PACT system from single finger to larger complex anatomical structures like the human hand.

This approach integrates the benefits of both advanced hardware configuration and cutting-edge machine learning techniques, offering a comprehensive solution to the inherent challenges in PACT of complex and large-scale targets.

2. Methods

2.1. Development of a large field-of-view photoacoustic computed tomography system

To construct our advanced PACT system, we integrated eight P6–3 transducers (center frequency of 4.5 MHz and bandwidth of 66%) in a half-ring configuration (radius of 90 mm) to optimize the field of view and imaging quality (Fig. 1(a)). The transducers were connected to four 256-channel NXT data acquisition systems (Verasonics, USA) and an ultrasound (US) machine to capture comprehensive imaging data. The laser source (Phocus mobile, OPOTEK, USA) provides the pulsed laser for PA excitation (Supplementary Fig. S1). We use 1064-nm wavelength excitation and optical illumination was achieved using an 8-pole optical fiber (aperture size of 6 mm) setup designed to focus light uniformly onto the target area (Fig. 1(a)). Fig. 1(b) depicts the configuration geometry, illustrating the arrangement of the transducers and optical fibers around the imaging region.

Fig. 1. Performance of a large field-of-view photoacoustic computed tomography (PACT) system.

Fig. 1.

(a) System schematic of the developed photoacoustic computed tomography. (b) Schematics of 1-view and 8-view PACT configurations. (c) US B-mode image of a tungsten wire target. (d) Comparison of 1-view and 8-view PAUS images of the tungsten wire target. (e) Comparison of spatial resolutions in US and PA images between 1 and view and 8-view. FB, optical fiber; TR, ultrasound transducer; US, ultrasound; PA, photoacoustic.

To evaluate the system characteristics, particularly spatial resolution, we employed a phantom consisting of 50-μm tungsten wires positioned radially within 3% agar, which included 0.05% TiO2. This configuration is illustrated in Fig. 1(c), showing the 8-view US images. The uniform radial positioning of tungsten wires in the agar-based phantom serves as a structured test for detecting the spatial resolution improvements facilitated by the multi-view setup. Fig. 1(d) compares the PACT maximum projection amplitude (MAP) images corresponding to the number of detectors (views): 1-view, 2-view, and 8-view configurations. This comparison elucidates the advantages of using multiple views in capturing detailed and accurate images. To quantitatively assess the spatial resolution improvements, we analyzed line profiles of one tungsten wire in Fig. 1(c–d). The full width at half maximum (FWHM) spatial resolution across different configurations for the x and z axes, respectively, are: US: 190 μm, 210 μm; 1-view PA: 1.09 mm, 469 μm; 2-view PA: 516 μm, 480 μm; 8-view PA: 176 μm, 352 μm. The results indicate that in the 8-view PA imaging setup, the lateral resolution significantly improved, approaching isotropic spatial resolution. These enhancements illustrate the capability of the developed system to maintain axial resolution while improving lateral resolution, which is a critical factor in achieving high-quality PA images. To further characterize the finite-aperture multi-array geometry, we evaluated field-dependent resolution by measuring the tungsten-wire point-spread function at increasing radial offsets. The lateral (X-axis) FWHM remained ~180–200 μm up to a 20-mm radius. At a 30-mm radius, the FWHM remained within ~200–300 μm for most directions, except for the outermost array direction (Dir. #1) with minimal multi-view contribution. Beyond 40 mm, the spatial resolution degraded substantially across all directions. Based on these trends, we define a practical achievable imaging radius of 30 mm (i.e., a 60-mm tomographic diameter) for the current illumination–detection configuration (Supplementary Fig. S2). We also validated the multi-array geometry and co-planarity by comparing the peak localization across individual arrays and the 8-view compounded reconstruction (Supplementary Fig. S3). The sub-millimeter dispersion of the wire peak confirms consistent array alignment and imaging of the same physical plane.

2.2. Deep learning strategy for enhancing image quality

To explore the in-vivo imaging capability of our developed system, we performed in-vivo human hand imaging. Fig. 2(a–b) illustrates the schematic setup for these two scenarios. In our PACT system, the 8-pole optical fibers are directed toward the center of the transducers’ half-ring shape. In single-finger imaging, this configuration ensures that the transducer detection plane and the laser illuminated region on the target are well-aligned (Fig. 2(a)). This alignment facilitates uniform illumination and detection, thus resulting in high-quality images. On the other hand, for larger targets such as a human hand, the imaging target must be positioned away from the central focus to encompass the entire region (Fig. 2(b)). Additionally, in the large-target imaging configuration, the optical fiber outputs also are positioned farther away from the center axis of the half-ring geometry compared with the single-finger configuration. While single-finger imaging places the optical fibers close to the center axis to maximize local illumination, imaging of larger targets such as the human hand requires the fibers to be retracted to secure a sufficiently large imaging region of interest (ROI). Increasing the source-to-target distance allows the laser beam to diverge, thereby expanding the illuminated area and providing more uniform coverage across the target. In contrast, when the fibers are placed too close to the target, the illumination becomes highly localized, limiting effective coverage and degrading image quality for large-area imaging. This deviation leads to two primary challenges: First, the illuminated region from the optical fibers does not align with the transducers’ detection plane, causing misalignment. Second, the complex and larger shape of the imaging target introduces optical shadowing. These combined factors lead to non-uniform illumination and degradation of image quality.

Fig. 2. Overall deep-learning strategy.

Fig. 2.

(a) Schematic of the PACT experimental setup for single-finger imaging. (b) Schematic of the PACT experimental setup for human hand imaging. (c) Deep learning strategy for enhancing PACT images. PACT, photoacoustic computed tomography.

To mitigate these issues of misalignment and optical shadowing in larger targets, we propose leveraging a DL enhancement strategy. This approach aims to correct the inconsistencies introduced by the misaligned illumination and detection planes and to compensate for the optical shadowing stemming from the complex target geometry. By applying DL techniques, we aim to significantly improve the overall target structures and signal-to-noise ratio (SNR). We utilized an encoder–decoder architecture to overcome the limitations posed by optical shadowing and geometric misalignment in large-target imaging. The detailed network configuration is summarized in Table 1.

Table 1.

Neural network architecture.

Layer Type Input Channels Output Channels

Encoder Layer 1 Conv2D + ReLU + BatchNorm 1 16
Encoder Layer 2 Conv2D + ReLU + BatchNorm 16 32
Encoder Layer 3 Conv2D + ReLU + BatchNorm 32 64
Encoder Layer 4 Conv2D + ReLU + BatchNorm 64 128
Bottleneck Layer Conv2D + ReLU + BatchNorm 128 256
Upconv Layer 4 ConvTranspose2D 256 128
Decoder Layer 4 Conv2D + ReLU + BatchNorm 256 (skip connection) 128
Upconv Layer 3 ConvTranspose2D 128 64
Decoder Layer 3 Conv2D + ReLU + BatchNorm 128 (skip connection) 64
Upconv Layer 2 ConvTranspose2D 64 32
Decoder Layer 2 Conv2D + ReLU + BatchNorm 64 (skip connection) 32
Upconv Layer 1 ConvTranspose2D 32 16
Decoder Layer 1 Conv2D + ReLU + BatchNorm 32 (skip connection) 16
Final Layer Conv2D 16 1

During the training phase, the input data (256 × 256 pixels) included PA images reconstructed from (i) adjacent-4, (ii) random-4 detector array selections, and (iii) element-level sparse sampling using only 128, 64, or 32 of the 1,024 total elements (1/8, 1/16, and 1/32 sparsity). Together, these conditions reproduce the limited angular coverage and reduced detection density characteristic of large-target PACT. As illustrated in Supplementary Fig. S4, the resulting degradations exhibit representative streak-like artifacts and fragmented or missing boundaries that closely resemble those observed in large-target hand imaging. The reference images were reconstructed using all eight detectors, providing high-quality reference images (Fig. 2(c)).

During the inference phase, the 8-view PA image of multiple fingers acquired from all eight detectors is applied to the trained network. To accommodate the large PA images, the inference process was conducted by dividing the input image into 256 × 256 pixels segments, with a 5-pixel overlap along both the width and height axes, and applying the model to each segment individually. These segments were then systematically stitched together through a sliding window approach to reconstruct the entire image. The results from this inference were then compounded to produce the final enhanced image.

For the DL environment, we utilized PyTorch version 3.10 in conjunction with an NVIDIA GeForce RTX 2070 GPU, enabled by CUDA version 12.1. The training parameters were set with a batch size of 32 and an epoch count of 200. Optimization was performed using the Adam optimizer, with a learning rate configured at 0.0001. The loss function employed was a composite of the L1 loss and the Structural Similarity Index Measure (SSIM) loss [11]. Normalization adjustments were performed, with the loss contributions weighted at a ratio of 2:8, respectively, for L1 loss to SSIM loss.

2.3. Data preparation

Training and validation data were acquired from one volunteer and comprised seven independent finger-scan acquisitions, totaling 1,348 reconstructed cross-sectional slices. For each slice, we generated seven subsampled inputs (details in Results section: Cluster #1–#3, Random four-array, and Sparse 1/8, 1/16, 1/32) from the same raw RF measurement and reconstructed them using the identical beamforming pipeline. Each subsampled reconstruction (input image) was paired with the full-aperture reconstruction using all eight arrays (reference image), yielding 9,436 paired examples. The paired examples were split into training and validation sets (80/20 hold-out) with a fixed random seed. Delay-and-sum beamforming was used for image reconstruction. Each slice was min–max normalized to [1]. During training, paired augmentation (random horizontal/vertical flips and random rotation up to 30°) was applied identically to input and reference images.

2.4. Deep learning quality metrics

The image quality of input and inferred images was evaluated by using peak signal-to-noise ratio (PSNR), SSIM, and root mean squared error (RMSE). The PSNR is a metric that estimates discrepancies between two images with respect to the peak signal amplitude of the inferred image:

PSNR=10log10peak_value2/MSE

where peak_value is taken from the range of the image data type and MSE is the mean squared error.

MSE=1N∑I−IRef.2,RMSE=MSE

where I refers to the input or inferred images and IRef. refers to the corresponding reference image. The SSIM is defined as

SSIM(x,y)=2μxμy+C12σxy+C2μx2+μy2+C1σx2+σy2+C2

where μx is the average of x , σx is the variance of x , and σxy is the covariance of x and y.

2.5. In-vivo imaging

Six healthy male volunteers were consented for the in-vivo imaging of the hand. All imaging procedures followed the protocol approved by the Stanford Institutional Review Board (protocol #58148). The laser fluence varied depending on the target area; from 30.4 mJ/cm2 to 10.2 mJ/cm2. In all cases, the laser fluence remained below the ANSI safety limit of 100 mJ/cm2 at a wavelength of 1064 nm. For 3D imaging, linear scanning was performed with a step size of 0.2 mm. The ROI for 2D single-finger imaging measured 25 mm × 25 mm with a resolution of 256 × 256 pixels and the ROI for hand imaging measured 100 mm × 100 mm with a resolution of 1024 × 1024 pixels.

3. Results

3.1. In-vivo PACT human hand imaging

In-vivo human single-finger imaging was performed using our developed PACT system. Fig. 3(a) presents the PA MAP image of a single finger, while Fig. 3(b) provides the PA depth-encoded image, revealing vessel structures from the skin surface. Notably, in Fig. 3(b), peeling approximately 2 mm from the skin surface reveals the proximal subungual arcade, and at a depth of roughly 7 mm, the proper digital artery becomes clearly discernible [34]. These results demonstrate the system’s capacity for high-quality anatomical visualization at substantial depths.

Fig. 3. In-vivo PACT human hand imaging.

Fig. 3.

(a) PA maximum amplitude projection (MAP) image of a single finger. (b) PA depth-encoded image of a single finger. (c) Comparison of PA cross-section images with varying numbers of detectors. (d) PAUS MAP image of a full hand. (e) PAUS cross-section image of multiple fingers.

Fig. 3(c) depicts the cross-sectional PACT image, where the yellow dotted line indicates the lateral profile and the red line denotes the axial profile used for quantitative evaluation. FWHM measurements were performed on a superficial vessel along these two profiles to assess how spatial resolution improves with an increasing number of detector views. The resulting FWHM values were as follows: 1-view; 1.37 mm, 0.52 mm; 2-view: 0.65 mm, 0.52 mm; 4-view: 0.44 mm, 0.51 mm; 8-view: 0.45 mm, 0.35 mm; lateral and axial, respectively. These measurements highlight the pronounced enhancement in lateral resolution achieved when multiple detector views are incorporated.

When imaging larger anatomical targets, such as a human hand including multiple fingers, several limitations became evident. Fig. 3(d–e) shows the US and PA images of the in-vivo human hand, illustrating the poor image quality in PA compared to single-finger imaging. Fig. 3(f–g) provides the cross-section US and PA images along the green dotted line in Fig. 3(d–e). The degradation in PA-image quality for the human hand is primarily induced by the lack of alignment between the optical path and the detection direction (Fig. 2(b)). Meanwhile, The B-mode US image clearly delineates bone structures, which exhibit high US reflectivity, as indicated by the yellow arrows in Fig. 3(f). Additionally, Supplementary Fig. S5 presents bone structures of the in-vivo human hand visualized by peeling off the skin signals. Additionally, the non-uniform optical distribution across the target leads to inadequate illumination coverage. This misalignment and non-uniform illumination are further compounded by the weak laser intensity, resulting in a low SNR and contrast-to-noise ratio (CNR). These factors collectively contribute to the poor image quality observed in human hand PACT imaging compared to the high-quality single-finger imaging.

3.2. Deep-learning enhanced in-vivo PACT single-finger imaging

To evaluate the performance of our network, we tested it using in-vivo single-finger PACT images (details provided in the Methods section). Fig. 4(a) illustrates the imaging configurations, where three different 4-view cluster conditions were generated: Cluster #1 using arrays #1–4, Cluster #2 using arrays #2–5, and Cluster #3 using arrays #3–6. The 8-view configuration, serving as the reference, incorporated signals from all eight transducer arrays to provide the highest-quality reconstruction. Fig. 4(b–c) present depth-encoded images acquired from each cluster condition at depths of 2.0 mm and 3.8 mm from the skin surface. Due to the limited angular coverage intrinsic to the 4-view clusters, regions outside the active detector span become invisible, and vessel structures appear fragmented or blurred, particularly along the missing-view directions. In contrast, the deep-learning–generated images exhibit markedly improved continuity and vascular visibility, closely approximating the reference images shown in Fig. 4(d). Fig. 4 (e–f) show cross-sectional images extracted along the white dotted lines in Fig. 4(b), comparing the input, DL-predicted, and reference reconstructions for each configuration. The DL-enhanced results successfully restore the skin surface boundaries (white arrows) and substantially reduce blurring of vascular structures (yellow arrows). These findings demonstrate that the network effectively compensates for the degraded image quality associated with 4-view clustered acquisitions, yielding outputs that closely resemble the 8-view reference reconstructions.

Fig. 4. Deep-learning enhanced in-vivo PACT single-finger imaging under the cluster condition.

Fig. 4.

(a) Schematic of the configuration conditions used for imaging: 4-view configurations (Cluster #1: arrays #1–4, Cluster #2: arrays #2–5, Cluster #3: arrays #3–6) compared to the 8-view reference. Depth-encoded images of the single finger from (b) 2.0 mm and (c) 3.8 mm based on the skin surface according to the configurations. (d) Depth-encoded images of reference. (e, f) Cross-sectional images along the white dotted line in (b) showcasing input, DL prediction, and reference images.

We further evaluated the network under the sparse configuration, which represents a more stringent limitation compared to the cluster condition. In this configuration, the PACT images were reconstructed using only a subset of the total 1,024 detection elements, corresponding to sparsity levels of 1/8 (128 elements), 1/16 (64 elements), and 1/32 (32 elements). Fig. 5(a–b) show depth-encoded images at 2.0 mm and 3.8 mm from the skin surface, respectively, for each sparsity level, with the 8-view reconstruction provided as the reference in Fig. 5(c). As the sparsity increases, the input images exhibit prominent noise and characteristic streak-shaped artifacts arising from insufficient detector sampling. Despite these severe degradation patterns, the DL-generated images display substantial recovery of anatomical structures. The network effectively suppresses streak artifacts, restores vessel continuity, and enhances the visibility of subdermal vascular features across all sparsity levels. Remarkably, even under the 1/32 sparsity condition, which provides only 32 active elements and produces extremely harsh input quality, the predicted images maintain strong similarity to the reference reconstructions. Fig. 5(d–e), which present cross-sectional views along the white dotted lines in Fig. 5(b), highlight this improvement: vessel boundaries that are nearly indistinguishable in the input images become clearly delineated after DL processing.

Fig. 5. Deep-learning enhanced in-vivo PACT single-finger imaging under the sparse condition.

Fig. 5.

Depth-encoded images of the single finger from (a) 2.0 mm and (b) 3.8 mm based on the skin surface according to each sparse configuration. (c) Depth-encoded images of reference. (d, e) Cross-sectional images along the white dotted line in (b) showcasing input, DL prediction, and reference images.

Quantitative evaluation using PSNR, SSIM, and RMSE further confirms these visual improvements (Table 2). Across all cluster and sparse conditions, the predicted images consistently outperform the corresponding inputs, demonstrating the robustness of the network in compensating for both limited detection coverage and severe element-level sparsity. In addition, we benchmarked the proposed network against modern alternatives, including DU-GAN [35] and CoreDiff [36], using the same dataset and evaluation protocol. As summarized in Supplementary Fig. S6, U-Net consistently achieved the highest PSNR and SSIM and the lowest RMSE across all conditions, with particularly strong gains in the most challenging sparse settings. These results indicate that a lightweight U-Net backbone provides reliable and high-quality restoration while remaining computationally practical for large-FOV inference.

Table 2.

Deep-learning metrics including PSNR, SSIM, and RMSE. PSNR, peak signal-to-noise ratio; SSIM, structure similarity index measure; RMSE, root mean square error.

Cluster #1
Input
Prediction Cluster #2
Input
Prediction Cluster #3
Input
Prediction

PSNR (dB) 27.6964 28.9823 28.7681 29.4420 28.0786 29.0172
SSIM 0.8010 0.8078 0.8060 0.8117 0.7743 0.7992
RMSE (×1e6) 0.043332 0.038194 0.039388 0.037024 0.042266 0.038571
Sparse 1/8
Input
Prediction Sparse 1/16
Input
Prediction Sparse 1/32
Input
Prediction
PSNR (dB) 30.3424 35.3976 22.7458 32.2991 19.4258 29.3405
SSIM 0.6677 0.8618 0.3771 0.8055 0.2278 0.7570
RMSE (×1e6) 0.034402 0.019754 0.076802 0.026905 0.109700 0.036293

We further evaluated the generalizability of our method by applying the network trained on a single volunteer to PACT data from five independent volunteers who were not included in training or model selection. Using the same frozen model and inference settings, the network consistently reconstructed anatomically coherent images across unseen subjects. As shown in Fig. 6(a), under the cluster configuration, the DL predictions exhibit improved vessel continuity and clearer skin-surface boundaries compared with the corresponding inputs. A similar trend is observed in the sparse configuration (Fig. 6(b–c)), where the network suppresses streak-shaped artifacts and restores vascular visibility despite severe input degradation.

Fig. 6. Generalizability of our deep-learning methodology.

Fig. 6.

Cross-sectional PACT single-finger images of a volunteer not included in the training dataset showcasing input and deep-learning prediction in (a) cluster configuration and (b) sparse configuration, and (c) reference images. (d) Subject-level quantitative performance across independent test subjects (N = 5), reported as mean ± standard deviation for PSNR, SSIM, and RMSE.

Quantitatively, we report subject-level performance by first computing per-subject mean metrics across slices and then summarizing results across subjects as mean ± standard deviation (Fig. 6(d)). In the sparse setting, the method achieved substantial improvements over the input baseline: for Sparse 1/16, PSNR increased from 23.36 ± 1.55 dB (input-to-reference) to 33.22 ± 1.87 dB (prediction-to-reference) (ΔPSNR = +9.86 ± 1.13 dB), SSIM increased from 0.437 ± 0.023 to 0.841 ± 0.026 (ΔSSIM = +0.404 ± 0.031), and RMSE decreased from 0.070 ± 0.012 to 0.022 ± 0.004 (ΔRMSE = −0.047 ± 0.009). For Sparse 1/32, PSNR increased from 18.98 ± 1.40 dB to 29.27 ± 1.78 dB (ΔPSNR = +10.29 ± 0.59 dB), SSIM increased from 0.238 ± 0.008 to 0.773 ± 0.029 (ΔSSIM = +0.535 ± 0.029), and RMSE decreased from 0.115 ± 0.017 to 0.035 ± 0.007 (ΔRMSE = −0.079 ± 0.011). In the cluster configurations, improvements were smaller but consistent across subjects (ΔPSNR = +0.68 to + 0.80 dB; ΔSSIM = +0.016 to + 0.024; ΔRMSE = −0.0027 to −0.0032). Collectively, these results demonstrate robust cross-subject generalization of the proposed DL-based enhancement framework.

3.3. Deep-learning enhanced in-vivo PACT human hand imaging

We further applied our DL enhancement strategy to in-vivo human hand imaging, which represents a more complex and optically heterogeneous target compared with a single finger. Using the full 8-view capability of our PACT system, we scanned a 70-mm segment of the hand and successfully acquired images of four fingers, from the index finger to the pinky finger. Fig. 7(a) presents depth-encoded images under three depth conditions: after peeling off 1.5 mm from the skin surface, after peeling off 3.0 mm, and across the 1.5–3.0 mm depth range. Consistent with the illumination geometry in Fig. 2(b), the pinky finger exhibits noticeably reduced image quality due to increased optical shadowing and its greater distance from the illumination fibers.

Fig. 7. Deep-learning enhanced in-vivo PACT human hand imaging.

Fig. 7.

(a) Depth-encoded images of the human hand at three depth conditions: (1) after peeling off 1.5 mm from the skin surface, (2) after peeling off 3.0 mm, and (3) corresponding to the 1.5–3.0 mm depth range. (b) DL-predicted depth-encoded images corresponding to the three depth conditions shown in (a). (c) Cross-sectional images extracted along the green and yellow dotted lines in (a). (d) DL-predicted cross-sectional images corresponding to those in (c). (e) Zoom-in images of the regions indicated by the yellow and green dotted boxes in (c), highlighting vascular structures and skin-surface boundaries.

To mitigate these degradation effects, we applied our network—trained solely on single-finger data—to the human-hand dataset. As shown in Fig. 7(b), the DL-predicted depth-encoded images exhibit clearer vessel continuity and improved structural visibility across all fingers. The cross-sectional comparisons in Fig. 7(c–d) further highlight the restoration of vascular boundaries and improved signal homogeneity. In the magnified regions (Fig. 7(e)), the network effectively recovers fine vascular structures (yellow arrows) and delineates the skin surface layers (white arrows), both of which are barely visible in the input images. We quantified the CNR between the blood-vessel signal and the artifact signal near the bone. The DL prediction achieved a substantially higher CNR (Prediction: 13.98) compared with the input (8-view: 8.76), demonstrating significant improvement in vessel visibility. Collectively, these results indicate that the proposed DL-enhanced PACT framework effectively compensates for variable CNR, optical shadowing, and heterogeneous imaging conditions, validating its applicability to complex anatomical targets such as the human hand.

4. Discussion and Conclusions

The results of this study demonstrate the significant potential of leveraging DL techniques to enhance PACT imaging, especially for complex and large-scale anatomical targets. Our developed PACT system, employing eight conventional linear arrays in a half-ring configuration, successfully achieved high image quality for single-finger imaging, but faced challenges during large-target imaging, such as multiple fingers. These challenges were primarily due to optical shadowing and misalignment between optical illumination and detection planes, leading to varied SNRs and poor visualization of vascular structures.

By integrating an encoder-decoder structure-based DL strategy, we were able to significantly mitigate these issues. The DL model, trained on single-finger PACT images, demonstrated its efficacy in restoring mutilated vessel and skin structural integrity, as well as improving overall image quality. This approach allowed for improving CNR and well-connected vascular structures across all fingers in human hand imaging scenarios. Furthermore, our DL strategy demonstrates high training efficiency and generalizability for human subjects. Although we trained the neural network using a dataset from a solely single human subject, the network could be effectively applied to five volunteer’s datasets. This capability arises from the fact that the overall anatomical structure, including vascular morphology, is generally similar across different human subjects. Consequently, the network can learn and accurately represent the anatomical structures of human hands based on PACT using only a single human dataset. Although GAN or diffusion-based restorers are powerful, our benchmarking shows that the encoder-decoder structure provides superior or comparable restoration quality with substantially lower deployment complexity. This DL strategy can be applied not only to the human hand but also to other limbs and organs (e.g., thyroid, breast), demonstrating the potential to universally provide high-quality PACT images even without the need for a premium PACT setup. We also emphasize the strength of this DL approach, which takes advantage of using a simple encoder-decoder architecture, rather than relying on burdensome and complex neural networks.

However, several limitations remain in our study. Firstly, although the DL enhancement demonstrated substantial improvements in finger and hand imaging, its generalizability to other anatomical structures was not investigated. Secondly, hardware constraints related to both the transducer array and optical delivery system inherently restrict the achievable image quality. Our system uses eight P6–3 phased-array probes arranged in a segmented half-ring geometry. The overall angular detection coverage is still limited by the finite aperture of each probe and the non-continuous angular sampling of the segmented half-ring, which can leave residual limited-view artifacts such as streaking and fragmented boundaries. Furthermore, the optical illumination provided by our 8-pole fiber bundle, with a small effective aperture of 6 mm, is not ideally matched to the acoustic detection aperture. Thus, uniform optical delivery is not always achieved even in single-finger imaging. Consequently, the 8-view compounded reconstruction used for supervision should be regarded as a best-available reference reconstruction under the current hardware constraints rather than an absolute ground truth of the initial pressure distribution. Lastly, although our analyses included both qualitative and selected quantitative metrics, more comprehensive evaluations, such as assessments of clinical relevance, diagnostic accuracy, and longitudinal consistency, will be required to establish the system’s reliability for broader biomedical applications.

Despite these limitations, in this work, we show that a simple and reproducible encoder–decoder network can effectively mitigate image degradation encountered in large-aperture PACT systems. Without relying on heavy or specialized network modules, the proposed model is trained using paired single-finger PACT data and subsequently applied, without retraining, to human-hand imaging. The results demonstrate that a network trained at a small anatomical scale can successfully generalize to larger and more complex targets, highlighting the feasibility of using DL to compensate for illumination- and geometry-induced degradation from optical shadowing and misalignment between source and detection in practical large-aperture PACT settings. Furthermore, while the current reference reconstructions are bounded by hardware and illumination constraints, the proposed learning-based enhancement consistently improves structural continuity under these practical limitations, and it is expected to integrate synergistically with future system upgrades toward higher-fidelity PACT imaging. The DL-based significant image-quality enhancements achieved suggest broader applicability to complex anatomical regions, potentially extending beyond PAD to other clinical domains requiring detailed vascular assessment.

Supplementary Material

1

Acknowledgements

This work was supported by the National Institute of Health grants R01CA271309, R01CA258807 and R01EB033967 (K. F.) and by the Ministry of Science and ICT grant RS-2024–00410679 through the National Research Foundation of Korea (S. C.).

Biographies

graphic file with name nihms-2164273-b0001.gif

Seongwook Choi is a postdoctoral scholar at Stanford University, Department of Radiology. He received his B.S. in Mechanical Engineering and Ph.D. in Convergence IT Engineering at Pohang University of Science and Technology (POSTECH), South Korea, in 2018 and 2023, respectively. Then, he worked in POSTECH Institute of Artificial Intelligence and joined Ferrara Lab at Stanford University in 2024. His research interests are the development of non-invasive biomedical imaging techniques including photoacoustic and ultrasound imaging and deep learning-based image processing.

graphic file with name nihms-2164273-b0002.gif

Katherine W. Ferrara is a Professor of Radiology at Stanford University, School of Medicine. She is a member of the National Academy of Engineering and a fellow of the IEEE, American Association for the Advancement of Science, the Biomedical Engineering Society, the Acoustical Society of America and the American Institute of Medical and Biological Engineering. Dr. Ferrara received her Ph.D. in 1989 from the University of California, Davis. Prior to her PhD, Dr. Ferrara was a project engineer for General Electric Medical Systems, involved in the development of early magnetic resonance imaging and ultrasound systems. Following an appointment as an Associate Professor in the Department of Biomedical Engineering at the University of Virginia, Charlottesville, Dr. Ferrara served as the founding chair of the Department of Biomedical Engineering at UC Davis. Her laboratory is known for early work in aspects of ultrasonics and has more recently expanded their focus to broadly investigate molecular imaging and drug delivery. Dr. Ferrara’s laboratory has received numerous awards including the Achievement Award from the IEEE Ultrasonics, Ferroelectrics and Frequency Control Society, which is the top honor of this society.

Appendix A. Supplementary data

Supplementary data to this article can be found online at https://doi.org/10.1016/j.ultras.2026.108081.

Footnotes

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

CRediT authorship contribution statement

Seongwook Choi: Writing – review & editing, Writing – original draft, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Katherine W. Ferrara: Writing – review & editing, Supervision, Resources, Methodology, Funding acquisition, Conceptualization.

Data availability

Data will be made available on request.

References

  • [1].Park J, Choi S, Knieling F, et al. , Clinical translation of photoacoustic imaging, Nat. Rev. Bioeng. 3 (2025) 193–212. [Google Scholar]
  • [2].Kim J, Lee J, Choi S, et al. , 3d multiparametric photoacoustic computed tomography of primary and metastatic tumors in living mice, ACS Nano 18 (2024) 18176–18190. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Kim J, Kweon JY, Choi S, et al. , Non-Invasive photoacoustic cerebrovascular monitoring of early-stage ischemic strokes In Vivo, Adv. Sci. 12 (2025) 2409361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Lee ES, Choi S, Lee J, et al. , Au/Fe/Au trilayer nanodiscs as theranostic agents for magnet-guided photothermal, chemodynamic therapy and ferroptosis with photoacoustic imaging, Chem. Eng. J. 505 (2025) 159137. [Google Scholar]
  • [5].Yang J, Choi S, Kim J, et al. , Multiplane spectroscopic whole-body photoacoustic computed tomography of small animals in vivo, Laser Photonics Rev. 19 (2025) 2400672. [Google Scholar]
  • [6].Wi J-S, Kim J, Kim MY, et al. , Theoretical and experimental comparison of the performance of gold, titanium, and platinum nanodiscs as contrast agents for photoacoustic imaging, RSC Adv. 13 (2023) 9441–9447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Choi S, Kim J, Jeon H, et al. , Photoacoustic computed tomography monitors cerebrospinal fluid dynamics and glymphatic function, Nat. Commun. (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Menozzi L, Li Z, Choi S, et al. , Light and metabolism: label-free optical imaging of metabolic activities in biological systems, Biomed. Opt. Express 16 (2025) 3770–3796. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Yang J, Choi S, Kim C, Practical review on photoacoustic computed tomography using curved ultrasound array transducer, Biomed. Eng. Lett. 12 (2022) 19–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Choi S, Kim J, Jeon H, et al. , Advancements in photoacoustic detection techniques for biomedical imaging, NPJ Acoustics 1 (2025) 1. [Google Scholar]
  • [11].Choi S, Yang J, Lee SY, et al. , Deep learning enhances multiparametric dynamic volumetric photoacoustic computed tomography in vivo (DL-PACT), Adv. Sci. 10 (2023) 2202089. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Kim M, Han JH, Ahn J, et al. , In vivo 3D photoacoustic and ultrasound analysis of hypopigmented skin lesions: a pilot study, Photoacoustics 43 (2025) 100705. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Gottschalk S, Degtyaruk O, Mc Larney B, et al. , Rapid volumetric optoacoustic imaging of neural dynamics across the mouse brain, Nat. Biomed. Eng. 3 (2019) 392–401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Nyayapathi N, Lim R, Zhang H, et al. , Dual scan mammoscope (DSM)—a new portable photoacoustic breast imaging system with scanning in craniocaudal plane, IEEE Trans. Biomed. Eng. 67 (2019) 1321–1327. [DOI] [PubMed] [Google Scholar]
  • [15].Lin L, Hu P, Shi J, et al. , Single-breath-hold photoacoustic computed tomography of the breast, Nat. Commun. 9 (2018) 2352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Li L, Zhu L, Ma C, et al. , Single-impulse panoramic photoacoustic computed tomography of small-animal whole-body dynamics at high spatiotemporal resolution, Nat. Biomed. Eng. 1 (2017) 0071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Nagae K, Asao Y, Sudo Y, et al. , Real-time 3D photoacoustic visualization system with a wide field of view for imaging human limbs, F1000Research 7 (2019) 1813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Lu N, Foiret J, Guo Y, et al. , Improving real-time ultrasound spine imaging with a large-aperture array, Sci. Adv. 11 (2025) eadw2601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Foiret J, Cai X, Bendjador H, et al. , Improving plane wave ultrasound imaging through real-time beamformation across multiple arrays, Sci. Rep. 12 (2022) 13386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [20].Choi W, Park E-Y, Jeon S, et al. , Three-dimensional multistructural quantitative photoacoustic and US imaging of human feet in vivo, Radiology 303 (2022) 467–473. [DOI] [PubMed] [Google Scholar]
  • [21].Chen T, Liu L, Ma X, et al. , Dedicated photoacoustic imaging instrument for human periphery blood vessels: a new paradigm for understanding the vascular health, IEEE Trans. Biomed. Eng. 69 (2021) 1093–1100. [DOI] [PubMed] [Google Scholar]
  • [22].Park E-Y, Cai X, Foiret J, et al. , Fast volumetric ultrasound facilitates high-resolution 3D mapping of tissue compartments, Sci. Adv. 9 (2023) eadg8176. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].Zhang Y, Wang L, Video-rate ring-array ultrasound and photoacoustic tomography, IEEE Trans. Med. Imaging 39 (2020) 4369–4375. [DOI] [PubMed] [Google Scholar]
  • [24].van Es P, Biswas SK, Moens HJB, et al. , Initial results of finger imaging using photoacoustic computed tomography, J. Biomed. Opt. 19 (2014) 060501. [DOI] [PubMed] [Google Scholar]
  • [25].Oeri M, Bost W, Sénégond N, et al. , Hybrid photoacoustic/ultrasound tomograph for real-time finger imaging, Ultrasound Med. Biol. 43 (2017) 2200–2212. [DOI] [PubMed] [Google Scholar]
  • [26].Agrawal S, Suresh T, Garikipati A, et al. , Modeling combined ultrasound and photoacoustic imaging: Simulations aiding device development and artificial intelligence, Photoacoustics 24 (2021) 100304. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Kruger RA, Kiser WL Jr, Reinecke DR, et al. , Thermoacoustic computed tomography using a conventional linear transducer array, Med. Phys. 30 (2003) 856–860. [DOI] [PubMed] [Google Scholar]
  • [28].Joseph Francis K, Boink YE, Dantuma M, et al. , Tomographic imaging with an ultrasound and LED-based photoacoustic system, Biomed. Opt. Express 11 (2020) 2152–2165. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Daoudi K, Van Den Berg P, Rabot O, et al. , Handheld probe integrating laser diode and ultrasound transducer array for ultrasound/photoacoustic dual modality imaging, Opt. Express 22 (2014) 26365–26374. [DOI] [PubMed] [Google Scholar]
  • [30].Yang J, Choi S, Kim J, et al. , Recent advances in deep-learning-enhanced photoacoustic imaging, Adv. Photonics Nexus 2 (2023) 054001. [Google Scholar]
  • [31].Li S, Chen Q, Kim C, et al. , Zero-Shot Artifact2Artifact: Self-incentive artifact removal for photoacoustic imaging, Photoacoustics 100723 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Jeong H, Oh S, Choi S, et al. , A hybrid diffusion model enhances multiparametric 3D photoacoustic computed tomography, Adv. Sci. (2025) e13624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Li S, Wang Y, Gao J, et al. , SlingBAG: point cloud-based iterative algorithm for large-scale 3D photoacoustic imaging, Nat. Commun. (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Cho S, Baik J, Managuli R, et al. , 3D PHOVIS: 3D photoacoustic visualization studio, Photoacoustics 18 (2020) 100168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Huang Z, Zhang J, Zhang Y, et al. , DU-GAN: Generative adversarial networks with dual-domain U-Net-based discriminators for low-dose CT denoising, IEEE Trans. Instrum. Meas. 71 (2021) 1–12. [Google Scholar]
  • [36].Gao Q, Li Z, Zhang J, et al. , CoreDiff: contextual error-modulated generalized diffusion model for low-dose CT denoising and generalization, IEEE Trans. Med. Imaging 43 (2023) 745–759. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

Data will be made available on request.

RESOURCES