Abstract
This study presents classification models trained to diagnose and grade prostate cancer using fresh prostate biopsies. We compare the performance of classification models with optimised sensitivity and specificity (standard models) with application-specific models designed to maximise sensitivity and negative predictive value (NPV). Standard models achieve 80% sensitivity and 81% specificity. Application-specific models, calibrated to 90% sensitivity and 95% NPV, are intended to provide clinicians with a tool they can use with confidence to support intraoperative decisions, specifically to improve tissue retention during biopsy procedures and to ensure clear surgical margins. To this end, we introduce a 5-layer algorithm that combines 5 application-specific models chosen for overall best performance. This algorithm can reduce the number of biopsy samples required for diagnosis by 47% while maintaining 90% sensitivity, 95% NPV, and 62% specificity. All models are independently validated using two large patient cohorts. These results support the targeted use of Raman spectroscopy for real-time tissue analysis in diagnostic and intraoperative settings. The technology’s clinical value as a decision-support tool aligns with the shared goal of pathologists and urologists to reduce the number of prostate biopsy cores while maintaining high sensitivity for clinically significant cancer. Prior studies have improved biopsy efficiency, but their performance has been variable, and concerns remain regarding underdetection of significant disease, revealing the need for approaches that improve biopsy efficiency without increasing diagnostic risk. The technology described here provides a realistic solution for targeted biopsy guidance to support more precise and evidence-based clinical decisions.
1. Introduction
Prostate cancer (PCa) is a significant healthcare problem worldwide. A recent American Cancer Society report showed that prostate cancer is the most diagnosed cancer in males in America, in Europe, in the central and southern African continents, and Australia, New Zealand, and Japan [1]. It is the second cause of cancer-related deaths in that group [1–5]. While the current burden of prostate cancer numbers is on Western countries, recent studies suggest that other parts of the world, such as Africa and Asia, are seeing the most significant rises in life-years lost [6,7]. From a healthcare infrastructure point of view, the global rise in cancer incidence has meant a considerable surge in biopsy samples requiring histology assessment in recent years. It will continue to do so in the following decades. Histology departments across the world are already reporting high levels of saturation, and an increasing burden on these resources is a problem that needs to be addressed [8,9].
Diagnosis of PCa involves a multi-stage process. Patients presenting symptoms and elevated PSA undergo MRI screening of the prostate region. If suspicious lesions are found, the surgeon will progress to a biopsy. For biopsy screening, clinicians typically use, including in this study, ultrasound-guided biopsies fused with pre-recorded MRI scans [10]. In some cases, however, simple ultrasound-guided biopsy protocols are used. In this study, biopsies were taken transrectally. The procedure often involves collecting 10-15 biopsy cores from 6-12 areas across the prostate with additional regions of interest identified on the MRI data [11–14]. The biopsies are then processed and examined by the histology specialists with an average turnaround time of 2-3 weeks [8,9]. The histology report can then be communicated to the clinicians and patients.
The biopsy procedure is invasive and can lead to side effects such as sepsis, haemorrhage, and urinary retention [15]. The large amount of tissue sent for assessment places a significant burden on the pathology labs. Furthermore, studies have shown that, on average, 50% of biopsy specimens contain only benign tissue, while up to 75% of the tissue mass extracted is benign or clinically non-significant [16]. Several potential solutions for improving this workflow are currently being explored. These include enhancing MRI image analysis and AI-driven histology slide analysis [17,18]. Indeed, MRI alone remains unable to diagnose PCa with the accuracy required [19]. While AI analysis of histology specimens has shown potential as a complementary tool to traditional histology, it does not obviate the need for pathology labs to process the biopsy cores, nor can it entirely replace human analysis by a trained pathologist [20,21]. Neither technique reduces the number of samples submitted for pathological evaluation; consequently, pathology labs must still process the same number of biopsy cores.
Prostate cancer treatment is traditionally through radical prostatectomy, radiation therapy or androgen deprivation therapy. Radiation therapy and androgen deprivation therapy require a positive prostate cancer diagnosis (for example, through biopsy) as well as prior to treatment. Consequently, a large proportion of patients will undergo radical prostatectomy. Urology surgeons face multiple challenges in the surgical management of PCa. Most of all, achieving maximal safe resection while preserving urinary and erectile functions by preserving the neurovascular bundle. These surgical challenges are exacerbated by the lack of intraoperative tissue assessment tools. Consequently, surgical decision-making relies largely on preoperative MRI, with intraoperative frozen histological sections used only in limited cases. Both approaches have substantial limitations that reduce diagnostic accuracy and disrupt surgical workflow. Although providing histological information during surgery, frozen section analysis offers limited value for real-time guidance. It typically requires at least 20 minutes for tissue preparation and assessment, delaying intraoperative decisions and extending anaesthesia time [22]. In addition, the rapid freezing required for this technique introduces artefacts that can compromise diagnostic reliability [23]. Finally, and most importantly, the technique requires immediate availability of a trained pathologist. This lack of in-surgery tissue analysis tools results in a high incidence rate of positive surgical margins (11−48%) [24–26].
A Raman spectroscopy instrument, used during biopsy collection or in-surgery, could provide real-time in vivo diagnostic information to inform clinicians. During biopsy collection, the technique can be used to reduce the number of biopsies taken, provide patients with a diagnosis on the same day, and reduce the downstream histology workload. Reducing the number of biopsies will also help reduce the procedure’s side effects, thereby improving the standard of care. Used during surgery, the instrument can enable real-time assessment of surgical margins, guiding the surgeon to remove all cancer while preserving the neurovascular bundle.
Recent research has shown that Raman spectroscopy has significant potential ex vivo as a research and diagnostic tool for cancer, with demonstrations reported for skin cancer [27–31], bone ageing [32], breast cancer [33–35], brain cancer (Raman sensing - [36–39])(multimodal sensing [40]), bladder cancer (multimodal sensing) [41], and prostate cancer [16,42–45]. Most studies to date have used only standard classification methods in isolated or ensemble models to obtain a binary output, typically separating benign and cancerous tissues [16,46–48].
In this report, we utilise a dataset of approximately 3300 Raman spectra obtained from fresh biopsy samples to demonstrate the development of several classification models designed for integration into a multi-layer classification algorithm. We demonstrate high classification performance for individual models and for a tailored multi-layer algorithm designed to provide real-time feedback to clinicians and minimise the number of required biopsies to diagnose PCa or for in-surgery guidance to ensure clear surgical margins. We focus our models’ performance on achieving high sensitivity (>90%) and high Negative Predictive Value (>95%) to provide clinicians with high confidence that the instrument will accurately detect the disease or that the tested tissue is benign and does not require biopsy or further excision. The study demonstrates the suitability of Raman spectroscopy for in-surgery use at the time of diagnosis or during prostatectomy procedures.
The clinical utility of the technology as a decision-support tool is underscored by the shared objective among pathologists and urologists to minimise the number of prostate biopsy cores while preserving high diagnostic sensitivity for clinically significant cancer. Multiple investigations have evaluated strategies to optimise biopsy yield by reducing the core number. Risk-stratification approaches have been proposed to individualise the number of biopsy cores based on patient-specific factors, demonstrating potential to decrease sampling intensity in higher-risk cohorts without compromising detection accuracy [49]. Other studies have explored pre-biopsy multiparametric MRI integration with biomarker and clinical data [50,51], using these multimodal models to guide targeted sampling. Although such combined approaches have shown variable success, concerns persist regarding possible underdetection of clinically significant lesions [51,52]. Collectively, this body of work highlights the need to enhance the efficiency of biopsy procedures while mitigating diagnostic risk. The technology described herein offers a potential real-time solution by enabling MRI-informed and MRI-independent targeted biopsy guidance, thereby augmenting diagnostic precision and supporting evidence-based clinical decision-making.
2. Methods
2.1. Study design
Our study was embedded within the surgical workflow to minimise disruption and maximise the number of patients included. Patients were seen by the specialist as part of their standard treatment. No patient selection was done for this study. Consent was obtained after the patient agreed to undergo the biopsy procedure. Patients presenting symptoms and elevated PSA received an MRI to image their prostate. If one or more regions of significant concern are identified, the patient proceeds to a biopsy. For this study, two protocols were used for guided biopsy collection. The clinicians either used ultrasound-guided (TRUS) protocols with a 6-region prostate mapping strategy, or a data-fusion system (Artemis) that enables overlap of MRI and live ultrasound, with a 12-region standard prostate mapping strategy + case-specific Region(s) of Interest (ROI). Raman spectra were collected from the cores before they were fixed in formalin and sent to pathology. The biopsy cores were analysed within minutes of excision and maintained hydrated with saline solution. Our goal was to measure as many biopsy cores as possible while limiting the time they spent in saline before fixation in formalin for pathology assessment. Thus, a single spot was measured on each biopsy core using a Raman spectrometer (10s integration time) coupled to a fibre optic probe. The clinician used an 18-gauge biopsy tool. A typical biopsy core was a 1 mm-wide cylinder about 15 to 18 mm long. The measured location, a 2 mm long spot, was marked using tissue dye before being sent to histopathology following standard protocols. Histologists reported on the entire core and the exact location of the measurement, which was labelled with tissue ink. Further details on patient consent, specimen handling and analysis, instrumentation, and pathology reporting can be found in [16]. We used a custom Raman spectrometer and probe from EmVision LLC. The probe has an external diameter of 1.65 mm and is coupled to a 785 nm wavelength-stabilised laser from Innovative Photonic Solutions. The probe consists of one excitation fibre surrounded by 12 collection fibres . The probe delivered 20 mW of laser light at its tip. The probe was lowered to approximately 1 mm above the specimen using a micrometre precision translation stage. The area of the tissue measured by the probe is approximately . The spectrometer used a back-illuminated deep-depletion CCD camera (iVac; Andor Technology). In total, the present study comprised 152 patients; 46 underwent TRUS biopsy, while 106 underwent 3D semi-robotic MRI and TRUS fusion (Artemis, Innomedicus) guided biopsy. Overall, 836 biopsy samples were extracted from 152 patients. We performed four repeat measurements (with identical acquisition parameters) at one location on each biopsy core. This enabled us to acquire 3342 Raman spectra. We then combined the pathology results and our spectra to train multiple diagnosis models and validate them using a leave-one-patient-out method. Repeat measurements (the four spectra acquired at the same location on each biopsy core) were kept together. They are not split between training and validation sets. The average age of the patient cohort, PSA level, and other information are detailed in Table 1. The study was conducted in accordance with the Declaration of Helsinki. All procedures performed in this study, including biopsy extraction, were conducted in accordance with the standard of care, institutional surgical guidelines, and approved ethical protocols. All participants were adults who provided informed consent. This project has been approved by the New Zealand Health and Disability Ethics Committees (HDEC, Ethics Reference Number 20/NTB/308).
Table 1. Study statistics. MRI - magnetic resonance imaging, TRUS-transrectal ultrasound.
| No. of Patients | 152 |
|---|---|
| Patient average age (range) | 66 (48-85) years |
|
| |
| PSA level (range) | 4.5 - 120 ng/mL |
|
| |
| Ethnicity, No. | |
| NZ European | 93 |
| M¯ori/Pasifika | 31 |
| Other | 28 |
|
| |
| Procedure type | No of patients |
| 3D semi-robotic MRI and TRUS fusion guided | Sub-cohort 1 = 106 |
| TRUS guided | Sub-cohort 2 = 46 |
|
| |
| No. of specimens | 836 cores |
|
| |
| No. of spectra (4 replicas/site) | 3342 spectra |
|
| |
| No. of biopsies (No. of spectra) | |
| Benign | Sub-cohort 1 = 439 (1755) |
| Sub-cohort 2 = 211 (844) | |
| Malignant | Sub-cohort 1 = 125 (499) |
| Sub-cohort 2 = 61 (244) | |
|
| |
| Cancerous cores (No. of spectra) | |
| Gleason pattern - GP3 | Sub-cohort 1 = 91 (364) |
| Sub-cohort 2 = 34 (136) | |
| Gleason pattern - GP4 | Sub-cohort 1 = 34 (135) |
| Sub-cohort 2 = 27 (108) | |
2.2. Data processing
The Raman spectra were preprocessed before being used to build our classification models. We kept pre-processing to a minimum to avoid affecting Raman features. Pre-processing was performed using in-house MATLAB code that included Raman shift calibration, cosmic-ray removal, instrument response correction, baseline correction, and spectra normalisation [16]. Multivariate statistical analysis was performed using MATLAB and the PLS Toolbox (Eigenvector Research Inc.). We initially used Principal Component Analysis (PCA) as an investigative tool to understand the variance in our data. Alongside PCA, we used Partial Least Squares (PLS) regression to identify spectral biomarkers in the Raman data that were relevant to the target diagnosis (see Table 2). PCA was not included further in the construction of the final models. To build our classification models, the full spectra were first analysed using PLS, and the resulting PLS scores were then used as input features for the SVM classifiers. For each model, we selected a set of features based on the desired diagnostic target and varied the number of features to avoid overtraining. To determine the number of latent variables for our SVM model, we used prior PCA results, which showed that 5, 20, and 50 latent variables were sufficient to capture 86%, 97%, and 98% of the variance, respectively. Figure 1 shows the mean pre-processed spectra of benign and cancerous tissues alongside an example of PLS features selected for cancerous tissue spectra.
Table 2. Models’ Prediction Target.
| Model 1 | Benign VS GP ≥ 3 (variation 1) |
| Model 2 | GP ≤ 3 VS GP ≥ 4 (variation 1) |
| Model 3 | GP ≤ 3 VS GP ≥ 4 (variation 2) |
| Model 4 | Benign VS GP ≥ 3 (variation 2) |
| Model 5 | Benign VS GP ≥ 3 (variation 3) |
Fig. 1.
(a) Mean pre-processed Raman spectra of benign and cancerous prostate biopsy core measurements. (b-e) Example PLS features selected for cancer-related spectral differences, with selected biochemical peak assignments indicated.
3. Model’s diagnosis performance
In this section, we examine the capability of a multi-layer classification algorithm to predict the presence of prostate cancer and compare it to a more traditional approach that uses a single classification model. We present our strategy for developing, testing, and selecting multiple classification models, which will then be concatenated and used as part of a larger, multi-layer classification algorithm. With access to a large dataset, our goal is to train, test, and cross-validate a wide range of models (varying their input data, classification objectives, etc.) using sub-cohort 1 and then compare their performance during independent validation using data from a separate cohort, sub-cohort 2. The performance of individual models is presented in Section 3.1, and that of the multi-layer algorithm is described in Section 3.2. For the multi-layer algorithm, independent validation tests focused on the ability to classify benign tissue against cancerous tissue (Gleason Pattern (GP) ≥ 3). It is essential to note that some of the models highlighted in Section 3.1 may exhibit limited performance when used independently. However, they have been selected as they enable better performance of the multi-layer algorithm.
3.1. Single model classification
Partial least squares regression was first utilised to reduce the dimensionality of the data. We then used Support Vector Machine (SVM) with a Radial Basis Function (RBF) kernel to construct our classification algorithms. To test the robustness and consistency of our models, we used a leave-one-patient-out cross-validation method. Multiple models were created by varying three different sets of parameters: (i) Data Selection, (ii) Model prediction target (see Table 2), and (iii) Model parameters (see Fig. 2). Overall, we tested approximately 900 models using various combinations of Input data, Model objective, and Model Parameters. Varying Input data is described fully in section 3.1.1 and includes selecting a subset of the data to train the models, as well as varying the number of latent variables for PLS feature selection.
Fig. 2.
Models training variations matrix.
We present below the 5 models that performed well individually (see Fig. 2 and Table 2). These models were selected to achieve the best-performing multi-layer classification algorithm.
3.1.1. Data selection
Data selection involves changing the number of features selected by a PLS feature selection, subsampling the data (i.e., excluding spectral repeats or balancing the numbers of benign and cancerous spectra), and adding noise to the data (data augmentation). Data selection primarily focuses on balancing our datasets. Due to our measurement protocols, our dataset comprises significantly more benign tissue measurements than cancer measurements. For sub-cohort 1, the dataset comprises 439 benign and 125 cancerous sites, corresponding to 1,755 spectra of benign tissue, 364 spectra of Gleason pattern 3, and 135 spectra of Gleason pattern 4. For example, in Method 4, we limited the number of benign tissue spectra used in our training model to the same number as the cancerous spectra.
As described in section 2.2, PLS regression was used to obtain the largest PLS scores most relevant to prostate cancer tissue identification (see Fig. 1) and reduce the dimensionality of the data. The identified features, the largest PLS scores, are then used as inputs to the SVM classifier. This approach ensures that the SVM classification focuses only on the variance in the data that best correlates with the presence of cancer. Five, twenty, and fifty PLS features were trialled. Limiting the number of model dimensions reduced the risk of overtraining of the SVMs.
We tested data augmentation, specifically Gaussian noise, in some of the models. Adding 5% noise to the input data reduced the variability of the models between leave-one-out validations. Tests indicated that a Gaussian noise of approximately 5% applied to each of the training sets after PLS feature selection was sufficient to enhance model consistency without compromising the average predictive power of the individual models.
To improve the completeness of our search and the complementarity between models, we included pathological information from the biopsy core beyond the boundaries of the region assessed by the Raman instrument in models 1 and 4. For Models 2, 3, and 5, we defined the target classification class based on the pathology report at the exact measurement location (labelled as Spot ≥ 3, or Spot ≥ 4 in Fig. 2). In contrast, for Models 1 and 4, the target classification class was defined based on the pathology assessment of the entire biopsy core (labelled as Core ≥ 3 in Fig. 2). This change allowed us to test the models’ ability to predict the presence of cancerous tissue within a biopsy core from a single-point measurement, thereby implicitly incorporating information from the surrounding tissue.
3.1.2. Model objective (Table 2)
For reference, Gleason Pattern ≤ 2 is considered benign tissue; GP3 is a low aggressiveness cancer (clinical non-significant); and GP4 and 5 are aggressive cancers (clinically significant) [16]. Changing the model objectives means asking the model to predict different grades of cancer, such as clinically significant (GP ≥ 4) versus clinically non-significant cancer (GP ≤ 3) or simply identifying cancer (GP4 or GP3 vs benign).
3.1.3. Model parameters
Finally, the changes to the Model parameters included box constraints and an epsilon term. Box constraints limit the number of vectors that the SVM model creates to be no greater than the specified value, and epsilon is the half-width of the incentive band. The incentive band is the region in which there is a gradient of certainty from the prediction of output class 1 to the prediction of output class 2.
3.1.4. Model performance
The models were trained and cross-validated on sub-cohort 1 of 106 patients (a 3-D semi-robotic MRI and TRUS fusion-guided cohort) using a leave-one-patient-out cross-validation approach. We cycled through the 106 possible iterations to test the stability and robustness of each model and observed stable output performance. This consistent performance highlights the comprehensiveness of our dataset. Consequently, we retrained each model on the complete set of 106 patients before independent validation. Figure 2 shows the models’ parameters found to be optimal at the cross-validation stage, before they were used for independent validation.
Independent validation was performed using sub-cohort 2, comprising 46 patients (undergoing TRUS-guided biopsy). For sub-cohort 2, the dataset includes 183 benign and 58 cancerous sites corresponding to 844 spectra of benign tissue, 136 spectra of Gleason pattern 3, and 108 spectra of Gleason pattern 4.
Among the models created, the 5 models described in Fig. 2 performed well individually. The classification performance of each model is illustrated in Fig. 3. The SVM classification models output a value between -1 and 1, with -1 indicating a "negative" output and 1 indicating a "positive" output. We used a simple threshold boundary to transform this regression output into binary outputs. Adjusting this threshold allows the model to be weighted towards maximising sensitivity or specificity.
Fig. 3.
(a-e) ROC curves for Models 1 to 5: Training used sub-cohort 1 (Blue), independent validation used sub-cohort 2 (Red).
Models 1 to 5 have an AUC of ROC of 0.86, 0.84, 0.86, 0.85, and 0.87, respectively. The validated models achieve sensitivity and specificity when optimising for both metrics, with Model 5 reaching 80% sensitivity and 81% specificity. However, in clinical cancer detection, high sensitivity and high NPV are often preferred. In the following, we therefore estimate the models’ performance for a set sensitivity of 90%. Models 1, 4, and 5, which have the same prediction target (benign VS ), have corresponding specificities of 45%, 53%, and 55%, respectively. Models 2 and 3 are designed to separate from . This classification is more challenging because both the positive and negative groups contain cancerous tissue. They have specificities of 50% and 35%, respectively, with a sensitivity of 90%. We include these two models to enhance the multi-layer algorithm’s ability to differentiate between clinically significant and clinically non-significant tissue.
3.2. Multi-layer classification algorithm
In this section, we aim to develop a multi-layer classification algorithm for intraoperative use that reliably identifies benign tissue with high negative predictive value (NPV) while minimising missed cancers through high sensitivity. The algorithm is optimised to support real-time tissue analysis and diagnostic decision-making. It is designed to maximise tissue retention during biopsy procedures (in other words, minimise the number of biopsies required to diagnose the disease) or to improve surgical outcomes by ensuring clear surgical margins (i.e., all cancer removed).
3.2.1. Algorithm prediction target
The classification objective of the multi-layer algorithm is to discriminate between benign tissue (Gleason pattern ≤ 2) and cancerous tissue (Gleason pattern ≥ 3). From a clinical perspective, when used as a guidance tool during biopsy procedures, tissue classified as benign would not be biopsied, whereas tissue classified as positive (cancerous) would be biopsied. As shown hereafter, this strategy can help reduce the total number of excised biopsy cores without compromising diagnostic outcome.
3.2.2. Algorithm structure
The models presented in the previous section were selected based on their performance when combined into a larger, layered algorithm. The number of layers in the multi-layer algorithm was determined empirically. A 5-layer algorithm was found to provide the best improvement with the fewest layers. A 6th layer added minimal predictive power while increasing the risk of overtraining, whereas the best combination of 4 models performed worse than the 5 selected models.
To build the multi-layer algorithm and optimise its performance, we tune the sensitivity threshold of each constituent model to identify benign samples with the highest possible negative predictive value. To assess improvements in sensitivity and negative predictive value with the addition of layers, the sensitivity and NPV at a given layer (e.g., layer n) are defined with respect to the cumulative model performance from layer 1 through layer n. Accordingly, spectra classified as positive (cancerous, GP ≥ 3) at a given layer are passed to the subsequent layer for further classification. In contrast, spectra classified as negative (benign) by any preceding layer are not reclassified but are retained in the benign output class when computing sensitivity and NPV at the nth layer.
We set the sensitivity to decrease by 2% between layers. Model 1 removed benign samples while maintaining a sensitivity of 98%. Model 2 removed benign samples while maintaining a sensitivity of 96%, and so on up to Model 5, which removed benign samples while maintaining a sensitivity of 90%. For each layer, we selected the best-performing model at the set sensitivity level. The corresponding NPVs from the 1st to 5th layer are 98%, 96%, 96%, 96%, and 95%, respectively. Sensitivity and NPV were selected as the primary performance metrics to reflect the standard surgical primary goal: complete tumour removal. High Sensitivity minimises false negatives (missed cancer), while high NPV ensures that tissue classified as benign can be safely left in the patient, prioritising oncological safety.
3.2.3. Algorithm performance
We assess the algorithm performance by varying the decision threshold and estimating, for each patient, the proportion of biopsy cores that can be confidently classified as benign and therefore not biopsied, i.e. ’saved’ compared to the total number of biopsies taken during the procedure under current clinical guidelines. This metric, referred to as the "Proportion of saved biopsies," is calculated as the ratio of a patient’s samples our algorithm classifies as benign to the total number of cores obtained under current standard-of-care biopsy guidelines and is used on the x-axes in Fig. 4. A value of 0% indicates that all cores are flagged as cancerous and that the same number of biopsies is sent to pathology as under current clinical practice. In contrast, 100% indicates that all assessed tissue is classified as benign and no biopsies are taken. Figure 4 illustrates the corresponding evolution of sensitivity and NPV.
Fig. 4.
Training (a), (c); Independent validation (b), (d). Sensitivity (blue) and Negative Predictive Value (NPV) (red) as a function of the Proportion of saved biopsies. (a)-(b) Model 5 testing and independent validation. The vertical lines highlight the number of saved biopsies at the set decision thresholds. (c)-(d) 5-layer algorithm testing, independent validation. The vertical lines highlight the number of saved biopsies at the set decision thresholds for each successive layer (from light grey to black).
We first report on the performance of the best single model, Model 5 reported above, to help calibrate the performance of the combined-models algorithm presented later in this section. The Proportion of saved biopsies with Model 5 is about 43%, as shown in Fig. 4 (a) and (b), with 90% sensitivity, 95% NPV, and 55% specificity. The next highest, Model 2, was 36%.
The multi-layer algorithm structure enables progressive removal of benign samples from the tested pool, removing benign samples with a very high confidence threshold in layer 1 (98% NPV), down to a confidence threshold of 95% in layer 5. This approach enables significant improvement of the specificity to 62% and also increases the Proportion of saved biopsies up to 47% while maintaining very high sensitivity and NPV thresholds at 90% and 95% respectively (Fig. 4 (c) and (d)).
4. Conclusion
The study demonstrates the capability of using a Raman spectroscopy instrument in the preoperative setting to identify likely malignant areas that can be targeted for biopsy, and intraoperatively to help guide the extent of surgical resection. We evaluate the performance of a custom-designed classification algorithm for a specific use-case scenario, aiming to confidently identify benign tissue in-surgery to guide surgical resection and thereby enable benign tissue retention.
We demonstrate robust classification models with stable performance across our large dataset. Standard models achieve 80% sensitivity and 81% specificity. Our application-specific models are configured to achieve a sensitivity of 90% and a negative predictive value of 95%, minimising the risk of missing cancers. We then test these models by estimating the proportion of biopsy cores that can be confidently classified as benign and therefore not biopsied, i.e. ’saved’. Individual classification models enable disease identification with 43% fewer biopsies while maintaining 90% sensitivity and 95% NPV. A dedicated 5-layer algorithm increases the Proportion of saved biopsies to 47% and the specificity by 7% up to 62% while maintaining very high sensitivity and NPV thresholds at 90% and 95% respectively.
The 152-patient cohort enabled the acquisition of an extensive dataset and the creation of statistically relevant classification algorithms. Models were independently trained and tested on two separate sub-cohorts (a training cohort of 106 patients and a validation cohort of 46 patients). Used during biopsy procedures, the system can help minimise the number of biopsies required to diagnose the disease, while during prostatectomy procedures, it can help ensure clear surgical margins (all cancer removed).
This study responds to the persistent need to improve the efficiency of prostate biopsy and prostatectomy procedures while maintaining reliable detection of clinically significant cancer. Existing approaches to reduce the number of biopsy cores, including risk stratification and multimodal MRI-based strategies, have shown mixed results and continue to raise concerns about missed disease. The findings presented here highlight the technique’s strong potential to accurately identify and grade prostate cancer using a handheld probe with short acquisition times (10 seconds). The combination of rapid measurement, real-time feedback to clinicians, and compact probe design supports its suitability for intraoperative use and readiness for clinical translation. It can enable targeted tissue excision guidance, applicable in both MRI-informed and MRI-independent settings, with the potential to enhance diagnostic accuracy while reducing unnecessary sampling. Both use-case scenarios would improve the current standard of care and demonstrate potential financial benefits for healthcare systems.
Acknowledgment
The authors would like to thank the Manukau SuperClinic (Urology Department).
Funding
Ministry of Business, Innovation and Employment https://ror.org/02jtq1b51 ( UOAX1806, UOAX2306); MedTech CoRE https://ror.org/02p521049 ( Rap Stage II - 2024); The Dodd-Walls Centre for Photonic and Quantum Technologies https://ror.org/05p61mv27.
Disclosures
The authors declare that they have no conflicts of interest. The authors used an algorithmic editing tool for grammar and style suggestions.
Author contributions statement. C.A and K.Z. conceived the experiments; M.J.D., H.L, and C.A. developed and tested the models; M.J.D and C.A. analysed the results. I.L. and M. L.C. performed the pathology assessments; A.A-L, M.R.P and K.Z. performed all clinical duties. A.P. provided clinical expertise. M.J.D. and C.A. wrote the manuscript. All authors reviewed the manuscript.
Data availability
The data underlying the results presented in this paper are not publicly available at this time; however, they may be obtained from the authors upon reasonable request.
References
- 1.McDowell S., “Cancer in men: Prostate cancer is number one for 118 countries globally,” Am. Cancer Soc. Res. News (2024). Accessed 11 July 2025.
- 2.International Agency for Research on Cancer , “Global cancer observatory,” World Health Organization. Accessed 11 July 2025.
- 3.Ferlay J., Ervik M., Lam F, et al. , “Global cancer observatory: Cancer today,” (2020).
- 4.Rawla P., “Epidemiology of prostate cancer,” World J. Oncol. 10(2), 63–89 (2019). 10.14740/wjon1191 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Culp M. B., Soerjomataram I., Efstathiou J. A, et al. , “Recent global patterns in prostate cancer incidence and mortality rates,” Eur. Urol. 77(1), 38–52 (2020). 10.1016/j.eururo.2019.08.005 [DOI] [PubMed] [Google Scholar]
- 6.Schafer E. J., Laversanne M., Sung H, et al. , “Recent patterns and trends in global prostate cancer incidence and mortality: An update,” Eur. Urol. 87(3), 302–313 (2025). 10.1016/j.eururo.2024.11.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Chu L. W., Ritchey J., Devesa S. S, et al. , “Prostate cancer incidence rates in africa,” Prostate Cancer 2011, 1–6 (2011). 10.1155/2011/947870 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Walsh E., Orsi N. M., “The current troubled state of the global pathology workforce: A concise review,” Diagn. Pathol. 19(1), 163 (2024). 10.1186/s13000-024-01590-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bychkov A., Schubert M., Constant Demand, Patchy Supply (The Pathologist, 2023). [Google Scholar]
- 10.Natarajan S., Marks L. S., Margolis D. J, et al. , “Clinical application of a 3d ultrasound-guided prostate biopsy system,” Urol. Oncol. 29(3), 334–342 (2011). 10.1016/j.urolonc.2011.02.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Park B. K., “Magnetic resonance imaging–guided prostate biopsy,” Cancers 13(22), 5647 (2021). 10.3390/cancers13225647 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Fütterer J. J., Briganti A., De Visschere P, et al. , “Can clinically significant prostate cancer be detected with multiparametric magnetic resonance imaging?” Eur. Urol. 68(6), 1045–1053 (2015). 10.1016/j.eururo.2015.01.013 [DOI] [PubMed] [Google Scholar]
- 13.Buller D. M., McLaughlin T., Staff I, et al. , “Outcomes of mri fusion-guided versus systematic standard prostate biopsies,” Can. J. Urol. 29, 10980–10985 (2022). [PubMed] [Google Scholar]
- 14.Pokorny M. R., de Rooij M., Duncan E, et al. , “Prospective study of diagnostic accuracy comparing prostate cancer detection by transrectal ultrasound–guided biopsy versus magnetic resonance (mr) imaging with subsequent mr-guided biopsy in men without previous prostate biopsies,” Eur. Urol. 66(1), 22–29 (2014). 10.1016/j.eururo.2014.03.002 [DOI] [PubMed] [Google Scholar]
- 15.Berry B., Parry M. G., Sujenthiran A, et al. , “Comparison of complications after transrectal and transperineal prostate biopsy: a national population-based study,” BJU Int. 126(1), 97–103 (2020). 10.1111/bju.15039 [DOI] [PubMed] [Google Scholar]
- 16.van Breugel S. J., Zargar-Shoshtari K., Aguergaray C, et al. , “Raman spectroscopy system for real-time diagnosis of clinically significant prostate cancer tissue,” J. Biophotonics 16(5), e202200334 (2023). 10.1002/jbio.202200334 [DOI] [PubMed] [Google Scholar]
- 17.Cai J. C., Nakai H., Kuanar S, et al. , “Fully automated deep learning model to detect clinically significant prostate cancer at mri,” Radiology 312(2), e232635 (2024). 10.1148/radiol.232635 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Saha S., Vignarajan J., Flesch A, et al. , “An artificial intelligent system for prostate cancer diagnosis in whole slide images,” J. Med. Syst. 48(1), 101 (2024). 10.1007/s10916-024-02118-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Boesen L., Nørgaard N., Løgager V, et al. , “Where do transrectal ultrasound- and magnetic resonance imaging-guided biopsies miss significant prostate cancer?” Urology 110, 154–160 (2017). 10.1016/j.urology.2017.08.028 [DOI] [PubMed] [Google Scholar]
- 20.Asadi-Aghbolaghi M., Darbandsari A., Zhang A, et al. , “Learning generalizable AI models for multi-center histopathology image classification,” npj Precis. Oncol. 8(1), 151 (2024). 10.1038/s41698-024-00652-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Darbandsari A., Farahani H., Asadi M, et al. , “AI-based histopathology image analysis reveals a distinct subset of endometrial cancers,” Nat. Commun. 15(1), 4973 (2024). 10.1038/s41467-024-49017-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Novis D. A., Zarbo R. J., “Interinstitutional comparison of frozen section turnaround time: A college of american pathologists q-probes study of 32,868 frozen sections in 700 hospitals,” Arch. Pathol. Lab. Med. 121, 559–567 (1997). [PubMed] [Google Scholar]
- 23.Amraei R., Moradi A., Zham H, et al. , “A comparison between the diagnostic accuracy of frozen section and permanent section analyses in central nervous system,” Asian Pac. J. Cancer Prev. 18, 659–666 (2017). 10.22034/APJCP.2017.18.3.659 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Silberstein J. L., Eastham J. A., “Significance and management of positive surgical margins at the time of radical prostatectomy,” Indian J. Urol. 30(4), 423–428 (2014). 10.4103/0970-1591.134240 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Beckmann K. R., O’Callaghan M. E., Vincent A. D, et al. , “Clinical outcomes for men with positive surgical margins after radical prostatectomy: Results from the south australian prostate cancer clinical outcomes collaborative community-based registry,” Asian J. Urol. 10(4), 502–511 (2023). 10.1016/j.ajur.2022.02.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.John A., Lim A., Catterwell R, et al. , “Length of positive surgical margins after radical prostatectomy: Does size matter? – a systematic review and meta-analysis,” Prostate Cancer Prostatic Dis. 26(4), 673–680 (2023). 10.1038/s41391-023-00654-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Bratchenko I. A., Bratchenko L. A., Khristoforova Y. A, et al. , “Classification of skin cancer using convolutional neural networks analysis of raman spectra,” Computer Methods and Programs in Biomedicine 219, 106755 (2022). 10.1016/j.cmpb.2022.106755 [DOI] [PubMed] [Google Scholar]
- 28.Boitor R. A., Varma S., Sharma A, et al. , “Diagnostic accuracy of autofluorescence-raman microspectroscopy for surgical margin assessment during mohs micrographic surgery of basal cell carcinoma,” Br. J. Dermatol. 191(3), 428–436 (2024). 10.1093/bjd/ljae196 [DOI] [PubMed] [Google Scholar]
- 29.Nieuwoudt M., Jarrett P., Matthews H, et al. , “Portable system for in-clinic differentiation of skin cancers from benign skin lesions and inflammatory dermatoses,” JID Innovations 4(1), 100238 (2024). 10.1016/j.xjidi.2023.100238 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Holtkamp H. U., Aguergaray C., Prangnell K, et al. , “Raman spectroscopy and mass spectrometry identifies a unique group of epidermal lipids in active discoid lupus erythematosus,” Sci. Rep. 13(1), 16452 (2023). 10.1038/s41598-023-43331-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Chung J. G., Holtkamp H., Nieuwoudt M, et al. , “The combination of raman spectroscopy and mass spectrometry to investigate cutaneous metallosis,” Br. J. Dermatol. 187(3), 447–448 (2022). 10.1111/bjd.20902 [DOI] [PubMed] [Google Scholar]
- 32.Nieuwoudt M. K., Shahlori R., Naot D, et al. , “Raman spectroscopy reveals age- and sex-related differences in cortical bone from people with osteoarthritis,” Sci. Rep. 10(1), 19443 (2020). 10.1038/s41598-020-76337-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Hanna K., Krzoska E., Shaaban A. M, et al. , “Raman spectroscopy: Current applications in breast cancer diagnosis, challenges and future prospects,” Br. J. Cancer 126(8), 1125–1139 (2022). 10.1038/s41416-021-01659-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Zhang Y., Li Z., Li Z, et al. , “Employing raman spectroscopy and machine learning for the identification of breast cancer,” Biol. Proced. Online 26(1), 28 (2024). 10.1186/s12575-024-00255-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Liu C.-H., Zhou Y., Sun Y, et al. , “Resonance raman and raman spectroscopy for breast cancer detection,” Technol. Cancer Res. Treat. 12(4), 371–382 (2013). 10.7785/tcrt.2012.500325 [DOI] [PubMed] [Google Scholar]
- 36.Desroches J., Jermyn M., Pinto M, et al. , “A new method using raman spectroscopy for in vivo targeted brain cancer tissue biopsy,” Sci. Rep. 8(1), 1792 (2018). 10.1038/s41598-018-20233-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Leblond F., Dallaire F., Ember K, et al. , “Quantitative assessment of the generalizability of a brain tumor Raman spectroscopy machine learning model to various tumor types including astrocytoma and oligodendroglioma,” J. Biomed. Opt. 30(01), 010501 (2025). 10.1117/1.JBO.30.1.010501 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Zhang L., Zhou Y., Wu B, et al. , “Intraoperative detection of human meningioma using a handheld visible resonance raman analyzer,” Lasers Med. Sci. 37(2), 1311–1319 (2022). 10.1007/s10103-021-03390-2 [DOI] [PubMed] [Google Scholar]
- 39.Zhang L., Zhou Y., Wu B, et al. , “A handheld visible resonance raman analyzer used in intraoperative detection of human glioma,” Cancers 15(6), 1752 (2023). 10.3390/cancers15061752 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Baria E., Giordano F., Guerrini R, et al. , “Dysplasia and tumor discrimination in brain tissues by combined fluorescence, raman, and diffuse reflectance spectroscopies,” Biomed. Opt. Express 14(3), 1256–1275 (2023). 10.1364/BOE.477035 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Baria E., Morselli S., Anand S, et al. , “Label-free grading and staging of urothelial carcinoma through multimodal fibre-probe spectroscopy,” J. Biophotonics 12(11), e201900087 (2019). 10.1002/jbio.201900087 [DOI] [PubMed] [Google Scholar]
- 42.Aubertin K., Desroches J., Jermyn M, et al. , “Combining high wavenumber and fingerprint raman spectroscopy for the detection of prostate cancer during radical prostatectomy,” Biomed. Opt. Express 9(9), 4294–4305 (2018). 10.1364/BOE.9.004294 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Pinto M., Zorn K. C., Tremblay J.-P, et al. , “Integration of a Raman spectroscopy system to a robotic-assisted surgical system for real-time tissue characterization during radical prostatectomy procedures,” J. Biomed. Opt. 24(02), 1 (2019). 10.1117/1.JBO.24.2.025001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Aubertin K., Trinh V. Q., Jermyn M, et al. , “Mesoscopic characterization of prostate cancer using raman spectroscopy: potential for diagnostics and therapeutics,” BJU Int. 122(2), 326–336 (2018). 10.1111/bju.14199 [DOI] [PubMed] [Google Scholar]
- 45.Picot F., Shams R., Dallaire F, et al. , “Image-guided Raman spectroscopy navigation system to improve transperineal prostate cancer detection. Part 1: Raman spectroscopy fiber-optics system and in situ tissue characterization,” J. Biomed. Opt. 27(09), 095003 (2022). 10.1117/1.JBO.27.9.095003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Lizio M. G., Boitor R., Notingher I., “Selective-sampling raman imaging techniques for ex vivo assessment of surgical margins in cancer surgery,” Analyst 146(12), 3799–3809 (2021). 10.1039/D1AN00296A [DOI] [PubMed] [Google Scholar]
- 47.Zú niga W. C., Jones V., Anderson S. M, et al. , “Raman spectroscopy for rapid evaluation of surgical margins during breast cancer lumpectomy,” Sci. Rep. 9(1), 14639 (2019). 10.1038/s41598-019-51112-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Gaba F., Tipping W. J., Salji M, et al. , “Raman spectroscopy in prostate cancer: Techniques, applications and advancements,” Cancers 14(6), 1535 (2022). 10.3390/cancers14061535 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Chen Z., Li Z., Dou R, et al. , “Personalized optimization of systematic prostate biopsy core number based on mpmri radiomics features: A large-sample retrospective analysis,” BMC Cancer 25(1), 116 (2025). 10.1186/s12885-024-13391-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Ippoliti S., Fletcher P., Orecchia L, et al. , “Optimal biopsy approach for detection of clinically significant prostate cancer,” Br. J. Radiol. 95(1131), 20210413 (2021). 10.1259/bjr.20210413 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Hagens M. J., Noordzij M. A., Mazel J. W, et al. , “Magnetic resonance imaging-directed targeted-plus-perilesional biopsy approach for prostate cancer diagnosis: “less is more”,” Eur. Urol. Open Sci. 43, 68–73 (2022). 10.1016/j.euros.2022.07.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Lee A. Y. M., Chen K., Tan Y. G, et al. , “Reducing the number of systematic biopsy cores in the era of MRI-targeted biopsy: Implications on clinically significant prostate cancer detection and relevance to focal therapy planning,” Prostate Cancer Prostatic Dis. 25(4), 720–726 (2022). 10.1038/s41391-021-00485-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data underlying the results presented in this paper are not publicly available at this time; however, they may be obtained from the authors upon reasonable request.




