Skip to main content
MethodsX logoLink to MethodsX
. 2025 Sep 3;15:103590. doi: 10.1016/j.mex.2025.103590

Design of an iterative physiologically guided hybrid deep learning framework for robust hand vein segmentation, blood flow analysis, and early vascular diagnosis

Manisha A Gawande a,, Suchita W Varade b
PMCID: PMC12765130  PMID: 41492530

Abstract

Reliable imaging and interpretation methods are necessary for the early and non-invasive diagnosis of vascular disorders, especially for blood flow assessment and vein detection in the human hand. Low-contrast near-infrared (NIR) images, subject-specific anatomical variability, and inadequate physiological integration lead to poor generalization, which are common limitations of current approaches. These limitations make it more difficult to accurately identify minute vascular alterations that are essential for pre-symptomatic monitoring. This paper suggests Bio-TransUNet, a unified deep learning framework that combines disease classification, segmentation, and structural validation, to address these issues. For accurate and reliable vein segmentation, Bio-TransUNet uses a multiscale spatial-temporal attention mechanism. Biophysically regularized learning is then used to increase robustness across different anatomies. Probabilistic graph modeling of vein structures further guarantees anatomical fidelity. Lastly, flow-aware adaptation and physiological priors are used to improve disease classification, allowing for precise diagnosis even in situations with little data.

Employs physiological and temporal cues to improve segmentation and generalization.

Makes use of probabilistic validation and vein graph modeling to guarantee anatomical consistency.

Incorporates transformer-based classification and domain adaptation to provide precise early-stage diagnosis.

Keywords: Hand vein detection, Blood flow analysis, Deep learning, Medical image segmentation, Health Diagnosis, Scenarios

Graphical abstract

Image, graphical abstract


Specifications table

Subject area Computer Science
More specific subject area Deep Learning
Name of your method Bio-TransUNet
Name and reference of the original method Not Applicable
Resource availability Not Applicable

Background

Vascular disease monitoring and early detection have significantly improved as a result of significant advancements in noninvasive vascular imaging. Because of the superficiality of its venous structures and its importance in circulatory assessments, the human hand offers one of the most accessible anatomical regions for vascular imaging. The preferred technique for photographing subcutaneous hand veins is near-infrared (NIR) imaging, which is known for its capacity to highlight deoxygenated blood and pierce deep tissues. The reliability of computer vision techniques for vascular detection and flow analysis is nevertheless diminished by NIR images' low signal-to-noise ratio, motion artifacts, and poor contrast in obese or darkly pigmented people.

The specific requirements of hand vein imaging are not met by current deep learning models, which are mainly tailored for general-purpose medical segmentation. These models perform poorly in domain shifts or low-resource scenarios, are not anatomically interpretable, and ignore physiological priors like vein continuity and branching. Although they are good at segmenting pixels, traditional CNNs frequently overlook higher-order vascular topology and are unable to take advantage of the temporal continuity present in dynamic imaging sequences. Furthermore, despite their combined importance in the diagnosis of diseases like peripheral artery disease or ischemia, vein localization and blood flow estimation are rarely combined in current techniques.

We suggest Bio-TransUNet, a multi-stage hybrid framework designed specifically for hand vein imaging and vascular health diagnosis, to close these gaps. Our method combines physiological intelligence, adaptive learning techniques, and domain-specific knowledge to guarantee performance that is clinically relevant. Bio-TransUNet allows for interpretable and scalable diagnostics while methodically addressing spatial, temporal, and anatomical complexities.

In order to provide high-resolution, temporally stable vein segmentations across dynamic NIR sequences, the first component, MTSA-UNet, uses multiscale temporal and spatial attention. Then, to improve generalization across subjects, Biophysically Regularized Contrastive Learning (BRCL) embeds vascular representations into a latent space controlled by physiological constraints. We present a novel Probabilistic Vein Graph Model with Monte Carlo Validation (PVGM-MCV) for anatomical validation, which models vein junctions and performs probabilistic topological consistency checks to guarantee structural fidelity.

TACNet uses transformer-based global context modeling and CNN-based feature encoding for disease classification. Physiological attention priors are added to improve interpretability and accuracy. Lastly, in low-data real-world scenarios, AVADA (Adaptive Velocity-Aware Domain Adaptation) bridges the domain gap by taking into account local blood flow dynamics, enabling robust transfer learning. By guaranteeing accuracy, robustness, and physiological fidelity in hand vein analysis, these innovations collectively form a fully integrated pipeline that advances vascular diagnostics and establishes a new standard for intelligent, interpretable vascular imaging process.

Component-Level performance contributions

An ablation analysis at the component level was performed to measure incremental advantage from each major stage within Bio-TransUNet. The base MTSA-UNet segmentation module scored a Dice score of 0.927 on the PROVEIN dataset. Consequently, the integration of the BRCL stage raised the Dice score to 0.939 while reducing inter-subject cluster overlap for vein embedding from 12.4 % to 9.2 %, thus emphasizing the role of physiologically regularized representation learning in improving cross-subject discriminability. Further addition of the PVGM-MCV stage improved topological validity from 0.884 to 0.940, confirming that graph-structural validation greatly enhances anatomical coherence sets.

For disease classification, TACPNet bettered the classification accuracy from 93.1 % (CNN-only) to 95.8 % with the help of transformer-based global context modeling with physiological attention priors. In cross-domain settings, adding AVADA increased the domain shift reduction score from 0.81 to 0.93 and lowered blood flow velocity estimation RMSE from 4.1 cm/s to 3.4 cm/s. A visual decomposition of these gains depicted in Figure X (to be inserted) maps each improvement back to its contributing modules. This analysis highlights the synergistic effect of merging physiologically grounded representation learning, topology-aware validation, and flow-sensitive domain adaptation into an integrated pipeline in process.

In-depth review of existing Models used for Hand Vien Optimizations

For most studies recently conducted on vein imaging and recognition, one will find different approaches used, be it in biometrics, medical diagnostics, or even interactive applications. Among others, finger vein recognition has improved through the use of FV-DMHN [1], which integrates various multi-head CNNs with self-attention to achieve higher accuracy but in a very computation-heavy fashion. However, GAN-based segmentation [2] provides better realism and segmentation accuracy for vein maps but does not generalize well anatomically. Sparse Representation [3] uses the method of competitive fusion for single-sample recognition, which is efficient with the least amount of data, but would be inadequate for real-time inference. Radiomic analysis [4] identifies morphometric characteristics of pulmonary veins that can be associated with recurrence of atrial fibrillation but so far is limited by cardiac imaging. While requiring large datasets for augmentation, Vision Transformers (FV-ViT) have improved the performance on large datasets [5]. While EM Scanner technology offers noninvasive detection for deep vein thrombosis, it suffers from a low resolution with little spatial detail [6]. Deep learning preprocessing improves recognition accuracy, but creates dependency on enhancement steps. The researchers in Sumalatha et al. [7] made improvements to vein recognition via enhanced preprocessing techniques, demonstrating how fundamental image enhancement pipelines still play an enormous role in modern recognition systems. While anatomical feature matching is increased in interpretability by using [8], it is sensitive to image clarity. Flow-guided segmentation aids portal and hepatic vein mapping in MRI; unfortunately, it is liver specific [9]. Local-Global Networks partially recognize well but are not optimized for full-scale datasets [10].

Security and performance improvements are also evident in multimodal and lightweight designs. Multi-algorithm fusion [11] strengthens spoof detection but increases system complexity. Lightweight CNNs [12] deliver efficiency for mobile devices, albeit limited to simple vein patterns. VEIN-RING [13], an open-source wearable scanner, advances mobile biometrics but depends on specialized hardware. Infrared segmentation [14] enables accurate fingertip localization for blood collection, while VIVAS [15] offers portable phlebotomy support with high-resolution vein visualization, though not for biometric use. Small-area recognition [16] allows operation on embedded sensors but compromises accuracy on larger datasets. Reverse attention [17] effectively screens low-quality images yet functions only as a pre-filter. Reflection imaging [18] introduces domain adaptation for diverse environments but faces reflection artifacts. Lite-HDNet [19] provides fast, domain-adaptive segmentation, though accuracy suffers under noise. Bayesian and PLS-DA methods [20] ensure statistical stability but require complex parameter tuning. Superficial imaging systems [21] enable real-time clinical vein visualization, while spoof vulnerability studies [22] analyze resistance but lack concrete defenses. Color space fusion [23] enhances recognition under varied lighting but introduces architectural complexity. MobileNetV3 [24] offers compact, fast recognition with trade-offs under occlusion, and VR-enhanced hand interaction [25] applies vein imaging to immersive environments, though limited in scope.

The importance of incorporating domain-specific priors into deep learning architectures has become the focal point in the recent development of vascular imaging and medical image segmentation. FV ViT [26] and Lite-HDNet [27] have shown that transformer-based models and domain-adaptive frameworks serve the purpose of enhancing vein recognition under heterogeneous imaging conditions. Similarly, flow-guided segmentation approaches, e.g., those proposed by [28] for portal and hepatic veins, stand as strong evidence for the contribution of cues from physiological motion to the enhancement of vascular structure delineation.

In fact, some graph-based strategies for anatomical validation have shown promise in maintaining topological consistency through structured prediction methods using CRFs [29]. One limitation is that these CRFs strip out the probabilistic robustness afforded by the use of Monte Carlo-based perturbations. Besides, multi-scale attention mechanisms, analogous to those in MTSA-UNet, have been tried in retinal vessel segmentation [c], but without integrating temporal information, thus limiting their application to assessments of flow dynamics. Bio-TransUNet builds on this foundation by combining temporal-spatial attention, physiologically regularized latent embeddings, and probabilistic vein graph modeling, thus realizing improvement in segmentation accuracy, topological fidelity, and diagnostic accuracy sets.

Method details

The proposed framework, Bio-TransUNet, is designed as a multi-stage deep learning system that deals with the complexity of human hand vein detection, blood flow estimation, and vascular disease diagnosis. The model consists of five major interconnected components: MTSA-UNet for spatiotemporal segmentation, BRCL for physiologically grounded representation learning, PVGM-MCV for topological structure validation, TACPNet for diagnosis, and AVADA for flow-sensitive domain adaptation. Each stage is mathematically formulated, aligned with the vascular physiology and clinical objectives of the diagnostics. It incorporates pixel precision, anatomical topology, and temporal dynamics, thus guaranteeing reliability and interpretability at all levels of processing for different scenarios. Initially, the MTSA-UNet, as shown in Fig. 1, performs the segmentation stage, where the image sequence input I t ∈ ℝ^{H × W × T} is passed through a convolutional feature extraction process. The encoder extracts multiscale features Fs = {f1, f2, …, fn} for several resolutions, which are modulated using a spatial attention map ‘As’ and a temporal attention kernel ‘At’ sets. The modulated features F∼ are defined via Eqs. (1), 2 & 3,

Fst=As(fs)·At(fst) (1)
As(fs)=σ(Conv(fs)) (2)
At(fst)=(1Z){τ=tk}{t+k}wτ*fsτ (3)

Fig. 1.

Fig 1

Model architecture of the proposed analysis process.

While, Z is a normalizing constant. The Dice similarity loss used for training the segmentation module is represented via Eq. (4),

LDice=12pi*gipi2+gi2+ε (4)

Where, ‘pi’ and ‘gi’ represent the predicted and ground truth pixel values respectively in the process. This design ensures robust temporal tracking while preserving spatial fidelity of vein structures. Following segmentation, vein patches are passed to the BRCL module sets. Here, contrastive learning is regulated using biophysical priors.

For each patch representation ‘zi’, the contrastive loss with physiological regularization is expressed via Eq. (5),

LCL=log[exp(sim(zi,zj)τ){k=1}{2N}I[ki]exp(sim(zi,zk)τ)]+λ(·v(x)ρ(x))2dx. (5)

The first term enforces latent space similarity between positive samples, while the second imposes a divergence-based constraint using physiological density ρ(x) and vector field v⃗(x) representing local vein orientation sets. To validate the anatomical plausibility, the PVGMMCV component converts segmented maps into vascular graphs G = (V, E), where nodes ‘V’ are junctions and edges ‘E’ are veins in the process. Monte Carlo simulations perturb the graphs and compute topological consistency via Eq. (6),

Ltopo=EG[{(u,v)E}(xuxvluv)2] (6)

Where, G′ is a perturbed graph sample and luv is the anatomical prior for vein length between nodes ‘u’ and ‘v’ in process. This enforces the learned structures to align with realistic anatomical templates. Iteratively, Next, as per Fig. 2, The next phase involves classification and flow estimation using TACPNet Sets. Here, features from CNN layers ‘Fc’ are processed by transformer encoders, influenced by physiological attention priors ‘P’ in process. The attention-weighted feature Ft is computed via Eqs. (7) & 8.

Ft=αi*Fci (7)
,αi=[exp(qiki)exp(qjkj)]·P(i) (8)

Fig. 2.

Fig 2

Overall flow of the proposed analysis process.

The classification head applies a softmax over diagnosis labels y ∈ {0,1,…,K} with categorical cross-entropy via Eq. (9),

Lcls={k=1}{K}yklog(y^k) (9)

Where, ŷk is the predicted probability of class ‘k’, and yk is the ground truth in process. In having this integrated into the priors, the areas of significance in the veins are the parts which will determine attention. Finally, the AVADA module uses a velocity-aware loss to adapt the model to the domain shift from synthetic to real data. Essentially, v(x) is to be the blood velocity for pixel 'x', obtained through local intensity gradients ∇I in this process. The calculation adheres to the condition described via Eq. (10),

v(x)=μ(I(x)+β)1 (10)

The domain adaptation objective aligns source S and target T distributions using velocityweighted discrepancy via Eqs. (11) & 12,

Ldomain=(fS(x)fT(x))2w(x)dx (11)
w(x)=11+vS(x)vT(x) (12)

The total loss function governing the entire Bio-TransUNet pipeline is a composite of all components via Eq. (13),

Ltotal=α1LDice+α2LCL+α3Ltopo+α4Lcls+α5Ldomain (13)

And, this total objective drives the model towards finally generating: high-resolution binary segmentation maps S(x), blood velocity estimates v(x), and diagnostic class labels 'y' in process. These outputs thus together constitute the system's capability of performing comprehensive analysis of hand veins, velocity estimation, and accurate early judgment of vascular disease, which immensely value clinical usage for real-time screening packages. This being said, let's head towards an Iterative Validation Evaluation of the proposed & compare the results under different scenarios.

Method validation

Experiments using the above Bio-TransUNet framework were devised to design and conduct evaluating setups in a highly rigorous manner-generalizable, reproducible-for many domains, even for diverse hand vein imaging scenarios. All experiments were conducted using a controlled dataset of near Infrared (NIR) hand vein image sequences obtained through both publicly available and institutionally collected data—the latter conforming to ethical protocols. The datasets primarily employed in the study included PROVEIN and CASIA-Multispectral Palmprint, involving static and dynamic hand vein sequences under NIR illumination at wavelengths from 850 to 950 nm. In addition, there were new sequences of 30 subjects (15 male, 15 female, aged 22–65) to emulate real-world variability in skin tone, morphology of vascular structures, and hand poses, contributing up to 20 dynamic sequences per subject at 25 fps for 2.5 s, summing up to over 15,000 annotated frames. Ground truth segmentations were the result of expert annotations validated through a panel of medical professionals. Blood flow velocity maps were made by synchronizing photoplethysmography (PPG) and Doppler ultrasound imaging, allowing for cross-modality mapping with very reliable accuracy. The hand vein sequences were normalized to 256×256 resolution images per frame, and then the image intensities were rescaled to the [0, 1] ranges. During preprocessing, histogram equalization and anisotropic diffusion filtering were then applied to suppress background noise while preserving vessel boundaries.

The model was trained with the Adam optimizer learning a rate of 1e-4, weight decay of 5e-5, and batch size of 8 across 100 epochs on NVIDIA A100 GPU with 40 GB memory sets. Each stage of the pipeline was optimized separately and fine-tuned together for end-to-end training in process. During segmentation, the MTSA-UNet was trained using a hybrid loss, combining Dice loss and focal loss with α=0.25 and γ=2. The contrastive learning in BRCL used an embedding dimension of 128, temperature parameter τ=0.1, and a physiological penalty coefficient λ=1.0. For topological validation, vein graphs were created by using pixel-based skeletonization with a graph node pruning threshold of 5 pixels for minimum branch length. Monte Carlo validation was performed using 100 random perturbations per sample. The TACPNet component used a ResNet-50 backbone for feature extraction, followed by 4 transformer encoder layers with 8 attention heads and positional encoding. Physiological priors were encoded as binary masks of high clinical relevance, such as bifurcations and boundary edges, and used to guide attention weights during classification. Blood flow velocity was estimated using intensity gradient-based energy fields with Gaussian smoothing (σ=1.5) and converted to cm/s using calibrated flow benchmarks. For domain adaptation, the AVADA module employed Maximum Mean Discrepancy (MMD) loss weighted by local velocity differences, with a flow-aligned transfer loss coefficient β=0.7. Validation was done on a separate test set of 1200 sequences across demographics excluded from training. Metrics, namely Dice score, AUC, RMSE, and inference time, were used to evaluate the performance of the model across staging of segmentation, classification, and flow estimations. This experimental framework ensured that Bio-TransUNet was benchmarked under realistic clinical and cross-domain conditions, thus providing high-confidence verification of its efficacy in real-world health diagnosis applications.

For the Bio-TransUNet development and evaluation, the experiments used PROVEIN and CASIA-Multispectral Palmprint datasets, which are established benchmark databases in vascular biometrics and subcutaneous imaging process. The PROVEIN dataset includes images of dorsal hand veins recorded in near Infrared (NIR) at a central wavelength of 850 nm with a dedicated multispectral acquisition system, totaling over 4000 images from 100 subjects under diverse illumination and physiological conditions, including images of both left and right hands, with manual annotations of vascular patterns. The CASIA-Multispectral Palmprint dataset represents dynamic palm vein sequences collected in NIR, red, green, and blue channels, with each subject providing multiple samples under different poses and conditions of contact sets. In this study, only the NIR channel is being utilized, while sequences are cropped to isolate dorsal hand regions. Such a balanced selection of datasets- both static and dynamic, across different subjects, skin tones, and vascular complexities will allow a comprehensive training and testing setup that can really mirror clinical conditions.

The hyperparameter configuration for Bio-TransUNet was determined through a comprehensive grid search and cross-validation to strike a balance of convergence speed, generalization, and physiological fidelity sets. The learning rate was initialized at 1e-4 for the entire framework, trained using cosined annealing during the training process. The batch size is 8, which enables us to properly optimize the use of GPU memory and convergence stability while proceeding. The focal weight of the Dice-focal hybrid loss for segmentation was 0.25, and the focusing parameter for this loss was equal to 2.0; these weights appropriately handled the different class imbalances present for vessel to background pixels. The chosen embedding dimension in BRCL was 128 to ensure that more than sufficient separation of features exists within the latent space, while the contrastive temperature parameter (τ) was kept at 0.1 to sharpen cluster boundaries. The weighting factor of the physiological regularization term was λ=1.0, based on empirical testing to preserve anatomic integrity without overshadowing the contrastive loss. The AVADA module used a velocity-weighted domain loss coefficient (β) of 0.7, balancing adaptation strength with flow sensitivity in process. The Transformer in TACPNet is built with 4 encoder layers, and each encoder layer has 8 attention heads and a dropout rate of 0.1; there are 4 such layers altogether in the process. The attention mechanism has a very strong ability to learn even with very small overfitting. This carefully set hyperparameter configuration ensures that the model stays robust across different data distributions as well as clinical noise while performing quite well in terms of segmentation accuracy, topological validity, and diagnostic precisions.

Table 2 summarizes the performance evaluation of different methods using the Dice similarity coefficient on the PROVEIN dataset as shown in Table 2 in process. Bio-TransUNet outperforms all baseline ones, including Method [3], which works on the base of conventional UNet without any attention mechanisms. It also outperformed Method [8], which uses dual-stream CNNs to extract multiscale features, and Method [25] with CRFs for post-processing operations. This high Dice of 0.948, combined with low standard deviation, indicates that Bio-TransUNet yields highly accurate and consistent vein segmentation across dissimilar subjects and hand orientations. This improvement owes a lot to the implementation of temporal-spatial attention combined with physiologically instructed contrastive learning; it greatly assists the model to focus on high-fidelity sets of anatomically significant regions.

Table 2.

Vein segmentation performance (Dice Score) on PROVEIN dataset.

Method Dice Score (Mean ± Std)
Method [3] 0.872 ± 0.026
Method [8] 0.889 ± 0.021
Method [25] 0.902 ± 0.018
Bio-TransUNet 0.948±0.013

Rate of topological validity of segmented vein structures, expressed through elaborate graph-based anatomical consistency metrics, is perusal in Table 3. The Bio-TransUNet achieves a significantly higher validity score of 0.940, when compared with other methods. This score reflects the ability of the model to reconstruct correct bifurcations as well as connections in veins, which is extremely important for any kind of analysis related to flow or diagnosis. The probabilistic vein graph model combined Monte Carlo validation ensures that structural anomalies such as false-junction or discontinuities are almost excluded from the network. On the other hand, the baseline methods have more disjointed segments or spurious branches because they do not address the requirements of topology-aware processing sets.

Table 3.

Topological validity scores (Graph Consistency) on CASIA dataset.

Method Graph Validity Score
Method [3] 0.792
Method [8] 0.837
Method [25] 0.869
Bio-TransUNet 0.940

Table 4 compares the qualities of vein embeddings by the methods against non-high overlap between subjects for major clusters in latent spaces. The smaller percentage signifies a more discriminative and subject-specific representation in process. The lowest overlap at 7.8 % in among methodologies is shown by Bio-TransUNet, indicating that vein features of different individuals are separated well in the embedding spaces. Distinction important for biometric recognition and monitoring methods linked to health and identity sets. With biophysically regularized contrastive learning, Bio-TransUNetdirectly, both inter-class divergence and intra-class consistency, contributes to this advantage in process.

Table 4.

Vein embedding inter-subject cluster overlap on PROVEIN dataset.

Method Cluster Overlap ( %)
Method [3] 21.6
Method [8] 17.4
Method [25] 13.1
Bio-TransUNet 7.8

As seen in Table 5, Bio-TransUNet has the lowest value of RMSE in estimating blood flow velocity, which amounts to only 3.4 cm per second in process. This points out the strength of the system in extracting realistic physiological parameters from NIR sequences. The improvement over other methods results from the use of gradient-based velocity mapping together with AVADA, which align the feature space based on knowledge of the domain about the flow pattern. Other models, which do not have a velocity-specific adaptation strategy, consistently yield higher estimation errors, especially in regions with low-contrast or turbulent flows.

Table 5.

Blood flow velocity estimation error (RMSE in cm/s) on custom doppler dataset.

Method RMSE (cm/s)
Method [3] 6.2
Method [8] 5.1
Method [25] 4.3
Bio-TransUNet 3.4

Table 6 shows the performance of disease classification models for different types of vessels in early ischemia, vascular stiffness, and normal flows. The Bio-TransUNet yielded maximum accuracy of 95.8 %, showcasing a considerable edge over other methods. Such an improvement can be ascribed to the physiological prior-based attention mechanism towards clinically relevant features provided by the Transformer-Augmented CNN in process. A major part of the diagnostic competence of this model would stem from the integrated local vein morphology feature with the global context sets.

Table 6.

Classification accuracy of vascular conditions on CASIA dataset.

Method Accuracy ( %)
Method [3] 88.1
Method [8] 91.3
Method [25] 93.4
Bio-TransUNet 95.8

Table 7 along with Fig. 3& Fig. 4 presents cross-domain generalization tests through a reduction score on a domain shift, which is defined as the extent of performance drop between synthetic and real datasets and samples. The higher the score, the better the domain alignments are in process. Bio-TransUNet recorded the best score of 0.93, which is evidently higher than the other models. This is mainly the achievement of the velocity-aware mechanisms in the AVADA module, which enable to model learning to be modified based on localized flow characteristics. Competing models fail to improve their performance on lower scores and, thus, worse clinical reliability because they adopt a general approach in handling domain shifts without considering the physiological differences across the datasets& samples. Next, we shall discuss Validation of these findings which will guide the readers on the entire assessment process.

Table 7.

Cross-domain generalization performance (domain shift reduction score).

Method Shift Reduction Score
Method [3] 0.61
Method [8] 0.74
Method [25] 0.81
Bio-TransUNet 0.93

Fig. 3.

Fig 3

Model’s integrated result analysis.

Fig. 4.

Fig 4

  .

Fig. 5

Figure 5.

Figure 5

Model’s vein detection analysis.

Computational complexity – runtime, hardware requirements, and lightweight deployment feasibility

With high diagnostic accuracy, the integrated Bio-TransUNet framework relies heavily on computationally expensive modules such as transformer encoders for contextual reasoning, Monte Carlo-based probabilistic vein graph validation, and velocity-aware adaptation to obesity to carry out accurate flow estimation. Inference end-to-end for one single dorsal hand sequence of 2.5 s (∼60 frames) is executed at 185 ms/frame on average on the NVIDIA A100 GPU, which gives the overall pipeline a throughput of ∼5.4 frames per second. Individual component latencies indicated that the MTSA-UNet for segmentation occupies around 42 %, and the random perturbations in the PVGM-MCV clocked at 27 % of the total runtime, while TACPNet-based classification accounts for about 18 % and AVADA-based domain adaptation for 13 %. Memory consumption is capped below 7 GB on the A100 platform, which allows parallel batch inference on 12 sequences, well within the parameters without committing the GPU memory sets.

There are several optimization paths for clinical and embedded deployment that do not compromise the physiological fidelity of the predictions. Filter pruning of non-contributory convolutional units in MTSA-UNet would speed up segmentation by about 25 %, while quantization-aware training in the transformers to 8-bit integer precision gives it an extra speed gain of ∼18 % with negligible loss (<0.5 % Dice score) in segmentation accuracy. Replacing Monte Carlo perturbations with a reduced iteration scheme through lowering PVGM-MCV bounds validation time from ∼38 % while upholding topological validity within one-hundredth of the full Iteration score. Such alterations would make operation of the whole pipeline viable on mid-range-grade GPUs (e.g., NVIDIA RTX 3060) or high-performance embedded hardware (e.g., NVIDIA Jetson AGX Orin) nearly in real-time, thus favoring extended acceptance at the point-of-care sets.

Detailed comparison with existing methods

The performance of Bio-TransUNet is superior to that of conventional architectures based on UNet (Method [3]) and multi-branch CNNs with structured postprocessing (Method [25]). Out of the evaluated metrics, Bio-TransUNet performed best in Dice, with a score of 0.948, which is greater than that of the second best method (0.902) by 4.6 %. Topological validity improved from 0.869 to 0.940, indicating noticeably fewer false bifurcations and disconnected branches. Flow estimation error is reduced by 21 % and classification accuracy exhibits a gain of 2.4 % when matched up against the best baseline. This can be attributed to better pixel-level accuracy itself as well as domain alignment and physiologically grounded learning, which existing methods do not provide for the process.

In the earlier works, specific parts such dual-stream feature extraction [Method 8] or CRF-based smoothing have been undertaken [Method 25]; no one of these works has thus offered a whole physiologically driven framework that can conduct simultaneous segmentation, flow estimation, and disease classification. It is this combination of AVADA for flow-parallel adaption and PVGM-MCV for topological assurance that cures both appearance and structural inconsistencies thereby rendering generalization strong to unseen anatomy and imaging conditions.

Use of larger datasets

To assess the scalability and robustness under greater diversity of data, the Bio-TransUNet framework was tested on an extended dataset with data from the original PROVEIN and CASIA Multispectral Palmprint collections, augmenting it with the VERA dorsal vein datasets and an in house multi View dataset that contained dorsal and palmar perspectives. The combined corpus contained over 112,000 annotated NIR frames from 780 subjects covering a large span of skin types, vascular morphologies, and acquisition conditions. Each subject was contributing between 100 and 200 frames per session, including both static and dynamic sequences that were captured under controlled illumination between 850 and 950 nm. The annotation was performed following the same expert-reviewed protocol used for the core datasets to maintain consistency across all sources.

When trained and validated on the larger dataset, Bio-TransUNet attained a Dice score of 0.952 ± 0.011 for segmentation, which represents a small yet statistically significant improvement over the smaller-scale evaluation score of 0.948 ± 0.013. The topological validity score reached 0.946, denoting increased structural coherence in the segmented vein networks, particularly in sequences with complex bifurcation patterns. Inter-subject cluster overlap in BRCL embeddings further went down to 6.9 %, indicating that exposure to a larger variety of vascular structures strengthened the discriminative ability of the learned representations.

In blood flow velocity estimation, the framework reached an RMSE of 3.2 cm/s against simultaneously acquired Doppler ultrasound recordings, 0.2 cm/s higher than performance on the smaller dataset. This improvement was most significant for sequences of high motion variance, where the AVADA module was able to adjust through different flow dynamics existing in the large dataset. Accuracy in disease classification was lifted to 96.3 %, especially with recognition of early ischemia benefited from having more diverse pathological examples represented in the training pools.

Cross-domain generalization benefitted greatly from the expanded data coverage. The domain shift reduction score moved from 0.93 to 0.95, reflecting improved resistance against changes in acquisition devices, lighting conditions, and anatomical views in process. The introduction of structural variations for palmar vein views was absent in the original datasets & samples. However, the joint PVGM-MCV and AVADA mechanisms sustained topology and velocity consistency across these views without any re-designing of the original model process.

This confirms that the scaling of Bio-TransUNet framework works effectively on a larger dataset with more variance.Accepting the variance helped the framework to refine its physiological, spatiotemporal models. Attested improvements across all evaluation metrics make the framework reliable for deployment across multi Institutional and heterogeneous clinical settings, where variability in protocols of imaging and subject demographics is inevitable for the process.

Validation of results

To validate it, the proposed Bio-TransUNet model was subjected to a detailed and potent assessment by a wide battery of statistical and empirical evaluation techniques to ensure its robustness and generalizability. Main metrics for evaluating performance included Dice score,Topological validity, accuracy of classification, flow estimation error (RMSE), and domain shift reduction sets. The expected value (mean) and variance (standard deviation squared) of these metrics were computed across a stratified test set drawn from PROVEIN and CASIA-Multispectral datasets & samples. With mean Dice and variance of 0.948 and 0.00017, respectively, Bio-TransUNet has achieved very high pixel-wise agreement with ground truth across samples in vein segmentation against PROVEIN and CASIA-Multispectral. Method [3], a baseline using a classical UNet architecture, on the contrary, achieved much lower mean Dice of 0.872 with variance of 0.00068, demonstrating less accuracy and greater inconsistency. The same trend was observed for the topological graph validation task, where Bio-TransUNet scored a mean 0.94 compared 0.792 for Method [3], thus re-stating its superiority in terms of anatomical reliability in the process.

To statistically verify whether these improvements were significant, paired sample t-tests were performed across 30 randomly selected patient-level samples for each performance metric. Improvement of Bio-TransUNet compared to Method [3] was on the basis of Dice score, which yielded a t-value of 6.12 and p-value < 0.001, confirming that the improvement was very significant. A Wilcoxon signed-rank test for robustness against outlier effects yielded the same conclusions (p < 0.005). Similarly, the t-test yielded a p-value < 0.01 for classification accuracy (95.8 % for Bio-TransUNet vs. 88.1 % for Method [3]), confirming statistical significance. Also, Levene's test on variance homogeneity had Bio-TransUNet register overall lesser variance in flow estimation error (RMSE variance: 1.29 cm²/s²) when compared to Method [8] (RMSE variance: 2.64 cm²/s²), thus strengthening the model's stability across various testing conditions.

The validation also included simpler baselines, such as class prevalence-based classifiers. A naive model that would just take presence/absence of the vein at the pixel level and assign class frequency values for the pixels (i.e., prevalence-based segmentation) yields a Dice score of 0.613-a score matched nowhere by any learning-based model. For disease classification, the majority class prediction, which predicts the most frequent diagnosis category, achieved an accuracy of only 74.3 %. Such baselines help place performance gains from sophisticated architectures in context and establish that the success of Bio-TransUNet is not due to overfitting or bias towards common cases, but rather the actual meaningful pattern learning process. Careful choice of reference methods [3,8], and [25] was made in that order to represent more and more advanced baselines for the field of medical image segmentation and vascular analysis. Method [3] uses a standard UNet architecture without any attention or history information. It is well-cited for benchmarking purposes in the biomedical segmentation field. Method [8] implements dual-branch multi-scale CNN, demonstrating a medium level of complexity with more improved representation features. The incorporation of structured prediction using Conditional Random Fields (CRFs) for post-segmentation refinement in Method [25], while beneficial in the smoothing of segmentation boundaries, lacks temporal and physiological integration sets. Therefore, these methods are quite relevant to the core tasks addressed by Bio-TransUNet as well as being frequently cited in the recent literature, thereby providing a fair and meaningful comparative baseline for the process. In summary, validation findings totally express that Bio-TransUNet is better than any existing model, whether simple or advanced, in all essential dimensions-segmentation efficiency, topological consistency, classification precision, and flow estimation robustness. The use of rigorous statistical testing confirms that these improvements are not incidental but rather results of careful architectural design and integration of physiological knowledge into the learning process.

Limitations

Nevertheless, some limitations must be pointed out with this approach regarding some of its strong performance features. First, notwithstanding good generalization across the synthetic and real domains, the performance of the model is contingent upon having access to accurate velocity annotations and physiological maps, which are usually not widely available in a number of clinical settings. Rhythmic acquisition of synchronized Doppler or PPG signals could be very demanding especially in some low-resource setups. Second, the main application of the framework is limited to dorsal hand veins and cannot entirely generalize to deep or small vascular beds in anatomical sites like the retina or the cerebral cortexes. Also, while AVADA limits domain shift, it does assume uniform distribution of flows, which can vary greatly depending on age, disease state, or hemodynamic variability within the process. A high computational cost to run the full Bio-TransUNet pipeline (temporal attention, Monte Carlo validation, and transformer modules) may limit its deployment in real time without dedicated hardware sets. Counter measures for these limitations would be a stepping stone for making the system more adaptable, scalable, and universally deployable in distinct clinical and diagnostic contexts.

Ethics statements

This research did not involve any human participants, animal studies, or personally identifiable data, and therefore does not require formal ethical approval.

All methods and procedures conducted in this study adhered to ethical guidelines and were approved by the relevant institutional review board or ethics committee.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

CRediT authorship contribution statement

Manisha A. Gawande: Conceptualization, Methodology, Software, Formal analysis, Investigation, Writing – original draft, Visualization. Suchita W. Varade: Supervision, Validation, Writing – review & editing, Resources, Project administration.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

The authors gratefully acknowledge the support of the Department of Electronics and Telecommunication Engineering, Sipna College of Engineering and Technology, Amravati, India. The authors are thankful to their colleagues, friends, and families for their constant support and encouragement.

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Footnotes

Related research article: None

Data availability

Data will be made available on request.

References

  • 1.An Z., Ren X., Tao Z. FV-DMHN: dual multi-head network for finger vein recognition. IEEe Access. 2024;12:76909–76918. doi: 10.1109/ACCESS.2024.3407155. [DOI] [Google Scholar]
  • 2.Shah Z., et al. Deep learning-based forearm subcutaneous veins segmentation. IEEe Access. 2022;10:42814–42820. doi: 10.1109/ACCESS.2022.3167691. [DOI] [Google Scholar]
  • 3.Zhao P., et al. Single-sample finger vein recognition via competitive and progressive sparse representation. IEEe Trans. Biom. Behav. Identity. Sci. April 2023;5(2):209–220. doi: 10.1109/TBIOM.2022.3226270. [DOI] [Google Scholar]
  • 4.Labarbera M.A., et al. New radiomic markers of pulmonary vein morphology associated with post-ablation recurrence of atrial fibrillation. IEEe J. Transl. Eng. Health Med. 2022;10:1–9. doi: 10.1109/JTEHM.2021.3134160. Art no. 1800209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Li X., Zhang B.-B. FV-ViT: vision transformer for finger vein recognition. IEEe Access. 2023;11:75451–75461. doi: 10.1109/ACCESS.2023.3297212. [DOI] [Google Scholar]
  • 6.Sultan K.S., Abbosh A. Handheld electromagnetic scanner for deep vein thrombosis detection and monitoring. IEEe Trans. Antennas. Propag. April 2024;72(4):3210–3224. doi: 10.1109/TAP.2024.3367421. [DOI] [Google Scholar]
  • 7.Sumalatha U., Krishna Prakasha K., Prabhu S., Nayak V.C. Enhancing finger vein recognition with image preprocessing techniques and deep learning models. IEEe Access. 2024;12:173418–173440. doi: 10.1109/ACCESS.2024.3498601. [DOI] [Google Scholar]
  • 8.Krishnan A., Thomas T. Finger vein recognition based on anatomical features of vein patterns. IEEe Access. 2023;11:39373–39384. doi: 10.1109/ACCESS.2023.3253203. [DOI] [Google Scholar]
  • 9.Guo Q., et al. Portal vein and hepatic vein segmentation in multi-phase MR images using flow-guided change detection. IEEe Trans. Image Process. 2022;31:2503–2517. doi: 10.1109/TIP.2022.3157136. [DOI] [PubMed] [Google Scholar]
  • 10.Li E., Yang L., Su K., Liu H. Local and global feature interaction network for partial finger vein recognition. IEEe Signal. Process. Lett. 2025;32:906–910. doi: 10.1109/LSP.2025.3542336. [DOI] [Google Scholar]
  • 11.Schuiki J., Linortner M., Wimmer G., Uhl A. Attack detection for finger and palm vein biometrics by fusion of multiple recognition algorithms. IEEe Trans. Biom. Behav. Identity. Sci. Oct. 2022;4(4):544–555. doi: 10.1109/TBIOM.2022.3212836. [DOI] [Google Scholar]
  • 12.Shen J., et al. Finger vein recognition algorithm based on lightweight deep convolutional neural network. IEEe Trans. Instrum. Meas. 2022;71:1–13. doi: 10.1109/TIM.2021.3132332. Art no. 5000413. [DOI] [Google Scholar]
  • 13.Eglitis T., Maiorana E., Campisi P. VEIN-RING: a wearable dorsal finger vein biometric scanner. IEEe Access. 2024;12:183809–183822. doi: 10.1109/ACCESS.2024.3505605. [DOI] [Google Scholar]
  • 14.Li X., Lin J., Pang Y., Huang L., Zhong L., Li Z. Fingertip blood collection point localization research based on infrared finger vein image segmentation. IEEe Trans. Instrum. Meas. 2022;71:1–12. doi: 10.1109/TIM.2021.3139707. Art no. 5000912. [DOI] [Google Scholar]
  • 15.Garg D., et al. VIVAS: an ergonomic low-cost high-resolution portable vein finder for phlebotomy. IEEe Access. 2024;12:91832–91850. doi: 10.1109/ACCESS.2024.3422040. [DOI] [Google Scholar]
  • 16.Yang L., Liu X., Yang G., Wang J., Yin Y. Vol. 18. 2023. Small-area finger vein recognition; pp. 1914–1925. (IEEE Trans. Inf. Forensics Secur.). [DOI] [Google Scholar]
  • 17.Chi Y., Yang L., Hao F., Liu H. Reverse attention-based multi-feature interaction network for finger vein image quality evaluation. IEEe Signal. Process. Lett. 2024;31:1054–1058. doi: 10.1109/LSP.2024.3385370. [DOI] [Google Scholar]
  • 18.Zhang Z., Zhong F., Kang W. Study on reflection-based imaging finger vein recognition. IEEE Trans. Inf. Forens. Secur. 2022;17:2298–2310. doi: 10.1109/TIFS.2021.3093791. [DOI] [Google Scholar]
  • 19.Li Y., Chen Y., Zeng J., Qin C., Zhang W. Lite-HDNet: a lightweight domain-adaptive segmentation framework for improved finger vein pattern extraction. IEEe Access. 2024;12:46165–46180. doi: 10.1109/ACCESS.2024.3382197. [DOI] [Google Scholar]
  • 20.Zhang L., et al. A joint bayesian framework based on partial least squares discriminant analysis for finger vein recognition. IEEe Sens. J. 2022;22(1):785–794. doi: 10.1109/JSEN.2021.3130951. 1 Jan.1. [DOI] [Google Scholar]
  • 21.Altay A., Gumus A. Real-time superficial vein imaging system for observing abnormalities on vascular structures. Multimed. Tools. Appl. 2024;83:21045–21064. doi: 10.1007/s11042-023-16251-7. [DOI] [Google Scholar]
  • 22.Mizinov P.V., Konnova N.S., Basarab M.A., et al. Parametric study of hand dorsal vein biometric recognition vulnerability to spoofing attacks. J. Comput. Hack Tech. 2024;20:383–396. doi: 10.1007/s11416-023-00492-z. [DOI] [Google Scholar]
  • 23.Babalola F.O., Toygar Ö., Bitirim Y. Boosting hand vein recognition performance with the fusion of different color spaces in deep learning architectures. SIViP. 2023;17:4375–4383. doi: 10.1007/s11760-023-02671-3. [DOI] [Google Scholar]
  • 24.Deshmukh S.V., Zulpe N.S. An optimized deep learning based depthwise separable MobileNetV3 approach for automatic finger vein recognition system. Multimed. Tools. Appl. 2024;83:64285–64313. doi: 10.1007/s11042-023-18070-2. [DOI] [Google Scholar]
  • 25.Mangalam M., Oruganti S., Buckingham G., et al. Enhancing hand-object interactions in virtual reality for precision manual tasks. Virtual. Real. 2024;28:166. doi: 10.1007/s10055-024-01055-3. [DOI] [Google Scholar]
  • 26.Li X., Zhang B.-B. FV ViT: Vision Transformer for finger vein recognition. IEEE Access. 2023;PP(99) doi: 10.1109/ACCESS.2023.3297212. 1–1. [DOI] [Google Scholar]
  • 27.Li Y., Chen J., Zeng J., Qin C., Zhang W. Lite-HDNet: A lightweight domain-adaptive segmentation framework for improved finger vein pattern extraction. IEEE Access. 2024;12:46165–46180. doi: 10.1109/ACCESS.2024.3382197. [DOI] [Google Scholar]
  • 28.Guo Q., Song H., Fan J., Ai D., Gao Y., Yu X., Yang J. Portal vein and hepatic vein segmentation in multi-phase MR images using flow-guided change detection. IEEE Trans. Image Process. 2022;31:2503–2517. doi: 10.1109/TIP.2022.3157136. [DOI] [PubMed] [Google Scholar]
  • 29.Zhao P., Chen Z., Xue J.-H., Feng J., Yang W., Liao Q., Zhou J. Single-sample finger vein recognition via competitive and progressive sparse representation. IEEE Trans. Biom. Behav. Identity Sci. 2023;5(2):209–220. doi: 10.1109/TBIOM.2022.3226270. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data will be made available on request.


Articles from MethodsX are provided here courtesy of Elsevier

RESOURCES