Abstract
Image-based machine learning tools are powerful resources for analyzing medical images, with deep learning-based semantic segmentation commonly utilized to enable the spatial quantification of structures visible in images. However, dataset generation and training of segmentation algorithms requires advanced programming skills and intricate workflows, limiting their accessibility to scientists without prior coding expertise. Here we present the step-by-step instructions to carry out automatic segmentation of medical images guided by a graphical user interface using the CODAvision algorithm. This workflow simplifies the process of semantic segmentation of microanatomical structures by enabling users to train highly customizable deep learning models without extensive coding expertise. The protocol outlines best practices for creating robust training datasets, configuring model parameters and optimizing performance across diverse biomedical image modalities. CODAvision enhances the usability of the CODA algorithm by streamlining parameter configuration, model training and performance evaluation, automatically generating quantitative results and comprehensive reports. We show the use of CODA to serial histology by demonstrating robust performance across numerous medical image modalities and diverse biological questions. We provide sample results in data types, including histology, magnetic resonance imaging and computed tomography. We demonstrate the diverse use of this tool in applications, including quantification of metastatic burden in in vivo models and deconvolution of spot-based spatial transcriptomics datasets. This protocol is designed for researchers with interest in rapid design of highly customizable semantic segmentation algorithms and a basic understanding of programming and anatomy.
Introduction
Deep learning has emerged as a tool for analyzing digitized biomedical images in various research applications1–11. Semantic segmentation models are commonly used as they can be quickly adapted to specific datasets for detecting structures across a variety of image types12–14. In addition, their relatively computationally lighter architectures (compared with some larger transformer and foundation models) make them amenable to researchers without access to high-performance computing. However, the implementation of these methods often requires advanced programming skills and intricate workflows, limiting their accessibility and widespread adoption outside of the computational biology community. In response, some popular deep learning workflows have been made more accessible through the development of graphical user interfaces (GUIs), which reduce or eliminate the need for extensive coding during implementation15–19.
To address this challenge, we developed a GUI aimed at simplifying the process of training highly customizable models for segmentation of biological structures in medical images. As a companion to this interface, we developed an extensive guide of best practices for annotation layer selection and annotation style for the construction of robust supervised models.
Overview of the procedure
The procedure is an extension of the CODA workflow20. CODA, a MATLAB-based pipeline for reconstruction of serial histological images into quantitative three-dimensional (3D) datasets, has been used extensively in biomedical research applications, including the study of pancreatic cancer progression, heart development, diabetic neuropathy and skin regeneration, among others20–40. CODA has shown technological power beyond its original presentation as a method to create 3D maps from serial hematoxylin and eosin (H&E)-stained histology41. Numerous recent studies have utilized one useful module of CODA, its image segmentation workflow, and integrated it with spatially resolved genomics, spatial transcriptomics and proteomics27,28,42, tissue stiffness29, antibody-based staining techniques30,31,43, organoid modeling32,44, and in vivo histology quantification34,40,45. We extract the segmentation module of the CODA package and dramatically improve its speed, usability and applicability to diverse biomedical image modalities. In the studies described earlier, the core implementation codes of CODA remained unchanged. The original implementation has three limitations that we address in the current protocol.
The original package is written as discrete functions in MATLAB. This language is not open source, and the format of the codebase made its implementation of the package challenging even for users with extensive programming experience. We address this through the translation of the CODA segmentation workflow to Python, optimization of the code speed and performance, and creation of a user-friendly GUI. We name this optimized workflow CODAvision. The original package lacked a description of the format and style of manual annotations required for training robust segmentation models. Recent groups have highlighted the importance of training dataset quality in deep learning approaches46,47. We address this through the generation of extensive user guides describing best practices for rapid generation of robust segmentation models. The original package demonstrated applicability to H&E images only. Here, we show how to use CODAvision for H&E samples, as well as other imaging types, including magnetic resonance imaging (MRI) and computed tomography.
In this workflow, we discuss the best practices for optimizing training datasets, constructing an intuitive user interface for parameter configuration and model architecture selection, and automatic generation of model performance reports and quantitative results for streamlined use in scientific experiments (Fig. 1).
Fig. 1 |. The CODAvision workflow.

A pipeline overview showing sequential steps: tissue annotation for dataset creation, GUI-guided parameterization, model training and quantitative analysis.
Comparison with alternative techniques
Compared with alternative techniques for quantifying structures in biomedical images described below, our workflow possesses differences that make it advantageous for certain research questions. In the proposed protocol, we guide users through the generation of robust datasets to train customizable deep learning models. Unlike most existing segmentation frameworks, CODAvision provides a user interface-guided workflow that allows manual annotations to be parametrized and automatically converted into optimized training tiles, enabling users with limited computational experience to efficiently generate trainable datasets. Another approach is use of large-scale foundation models and vision transformers that have been trained on millions of examples for object-detection tasks, or sometimes more specifically on extraction of meaningful features from biomedical images8,48–51. These methods differ from our protocol in that they are intended to be fully automated, where our workflow requires initial manual annotation.
For research projects where direct measure of a certain structure (such as cancer metastases in mouse histology, or something very specific such as a subtle phenotype of cancer cells that is not currently well defined) is desired, our workflow enables rapid segmentation of that specific anatomical structure. By contrast, for applications that benefit from a more holistic perspective—such as generalized feature extraction or the generation of attention maps for survival prediction—pretrained foundation models may be more suitable. Our workflow is designed to run on a standard desktop computer and requires only minimal programming expertise, making it broadly accessible. However, the deployment of large foundation or transformer models typically demands high-performance computing resources and advanced computational skills.
As a one-to-one comparison with widely used tools for biomedical image analysis, we evaluated our workflow against QuPath (version 0.5.1) pixel classification, nnU-Net and 3D slicer52 (Fig. 2). QuPath integrates annotation and classification within a single interface and provides real-time visual feedback for a smooth user experience. QuPath pixel classification is useful for rapid training of an algorithm to quantify positive signal in immunofluorescence data, for example. However, for more complex segmentation problems such as differentiation of subtle features in histology or radiology, more advanced algorithms are required. CODAvision uses semantic segmentation and a larger model architecture that is well-suited for complex classification tasks, although with the trade-off of relatively longer processing times (1–2 h) compared with QuPath (1–2 min). To quantitatively compare CODAvision segmentation to QuPath pixel classification, we generated a dataset of mouse lung histology. We annotated in Aperio ImageScope (using the overlapping annotation style suitable for CODAvision and described in detail in Procedure 2) as well as in QuPath (using non-overlapping annotations that may be more suitable for QuPath model training; Extended Data Fig. 1). In both annotation scenarios, CODAvision outperformed QuPath, achieving >90% overall accuracy compared with QuPath’s accuracy of ~30%. CODAvision segmentation also appeared more robust to correct handling of artefacts such as shadows and dust in the whole slide image compared with QuPath. Overall, while QuPath is a powerful tool capable of rapidly performing diverse image analysis tasks, external tools such as CODAvision offer advantages for certain segmentation of subtle anatomical structures.
Fig. 2 |. The classification of mouse lung histology using CODAvision, QuPath and nnU-Net.

A dataset of mouse lung histology was generated, and anatomical structures were annotated in these images using Aperio ImageScope: alveoli, bronchioles, stroma, cancer metastases, vasculature and background. One image was used for independent model testing, with all others used for training. CODAvision performed with >90% overall accuracy and >85% per-class precision and recall. QuPath pixel classification performed significantly worse, with ~33% overall accuracy and extensive over-calling of metastases in the testing image. nnU-Net was trained on the same masks generated through CODAvision, but the model struggled to generalize to WSI distributions where the object density and spatial context differ from the tiles used for training, showing significant misclassification across all tissue structures considered, with ~45% overall accuracy.
Furthermore, we evaluated the convolutional neural network (CNN) training of CODAvision to that of another commonly used segmentation algorithm, nnU-Net53 (Fig. 2). As the nnU-Net training pipeline does not include guidelines for annotation generation or training tile construction, we used the CODAvision-generated tiles to train a two-dimensional nnU-Net model (nnUNetTrainer with Dice–cross-entropy loss, deep supervision and fivefold cross-validation). Preprocessing followed nnU-Net’s built-in red, green, and blue (RGB) image functions. Training on the same NVIDIA RTX 5000 ADA graphics processing unit (GPU) required ~75 h for nnU-Net versus 25 min using CODAvision. The nnU-Net model showed substantial misclassification across tissue types, with overall accuracy on an independent test image of 45%, compared with 93.2% for the CODAvision-trained model.
Moreover, we compared CODAvision for MPRAGE (sagittal) brain MRI segmentation against 3D Slicer Segment Editor and the MONAI Bundle Zoo pretrained model (wholeBrainSeg_Large_UNEST_segmentation)54. 3D Slicer produced low-precision annotations with nonsmooth tissue borders, while the MONAI-based model misclassified anatomical structures in multiple regions (Extended Data Figs. 2 and 3). Although 3D Slicer is robust for organ-level 3D segmentation and visualization of medical images such as computed tomography or MRI, it is less suitable for multiple independent two-dimensional images such as histology. In cases where pretrained or intensity-based approaches fail to capture complex or subtle structures, CODAvision enables detailed custom segmentation.
Finally, popular software such as DeepCell and Cellpose are highly effective for nuclear segmentation in fluorescence or brightfield microscopy images. However, these models differ from CODAvision in that they are designed for instance segmentation, not customizable semantic segmentation of tissue components55,56.
Experimental design
The five main steps outlined in the protocol are: dataset creation, GUI-guided parameterization, model training, model optimization and image segmentation using a custom pretrained model.
Example datasets: to help users explore the CODAvision software and its features, we include a link to a sample dataset of annotated mouse lung histology. Results obtained from CODAvision analysis of these datasets are presented in the Anticipated results. We recommend that users initially run CODAvision on the supplied dataset and follow the detailed protocol described in the procedure section before analyzing new data.
Dataset creation
A dataset for this workflow consists of a set of biomedical images (for example, digitized histology, MRI or computed tomography) and their associated annotation metadata. The first step of the workflow is to identify structures in the dataset that can be distinguished from each other, and which of those are required for the research objective. An exhaustive list of structures is then made, and manual annotations of these structures are made in the freely available program Aperio ImageScope. The selection of these annotated structures directly corresponds to the research question the CODAvision analysis aims to address. In Procedure 1, we outline strategies for developing a robust training dataset that enables building custom segmentation models for the targeted structures.
GUI-guided parameterization
Once the manual annotations are complete, the next step in the CODAvision workflow is to define the model training parameters using a Python-based GUI. Users first install the CODAvision Python package, following the instructions for codes and dependencies outlined in Procedure 2. The GUI guides users through configuring settings for model training, including specifying the location of training and testing datasets, selecting an image resolution (downsampled files are generated using OpenSlide)57 and customizing model parameters. The choice of resolution is critical, balancing segmentation detail with computational efficiency. The GUI further enables the management of the manual annotated layers, enabling features such automatic removal of background pixels, combining or deleting annotation layers, and defining the nesting logic for managing overlapping annotations. Advanced settings enable further customization, such as adjusting tile size, batch size and model architecture (for example, DeepLabV3+ or UNet)12,58, allowing users to optimize training on the basis of their computational resources and dataset size. Once all parameters are set, the model will begin training.
Model training
The model training phase begins with tissue thresholding, guided by an interactive popup window that allows users to fine-tune the threshold cutoff for optimal tissue/background separation. Once the threshold is set, the preprocessing and model training proceed automatically. During this phase, the .xml annotation coordinate data will be imported and converted to .png annotation masks. These masks will be used to create training and validation tiles built using data augmentation techniques such as hue adjustment, scaling, rotation and Gaussian filtering to enhance dataset heterogeneity and model robustness. The chosen model architecture is trained on these augmented tiles. After training, the model performs inference on test images, generating a confusion matrix that showcases the precision, recall and overall accuracy. For each segmented image, the workflow will output a classified .tif mask and a colorized .jpg overlay, enabling users to rapidly review the model performance using both quantitative (confusion matrix) and qualitative (overlay images) review. The overall composition of each image will be saved in a .csv file, along with more detailed morphological calculations if desired by the user. The model results will be automatically summarized in a generated .pdf report. Users are encouraged to review these results to ensure the trained model meets the recommended performance benchmarks (for example, >90% overall accuracy and >85% per-class precision and recall, and visually acceptable results).
Model optimization
If the performance of the trained model is unsatisfactory, Procedure 4 outlines steps to efficiently retrain and optimize the model. Optimization strategies include review of the colorized mask overlay images to identify patterns of misclassification, determining whether any structures were missed during initial annotation, and adjusting model parameters to better align with the desired results. If misclassifications persist, it is recommended to expand the training dataset by adding new images rather than adding annotations to existing training images. These optimization steps can be repeated iteratively until satisfactory results are achieved, ensuring the model meets the recommended performance benchmarks.
Image segmentation with custom pretrained model
Once a model is trained, users can apply it to segment additional images beyond the original training set using a pretrained model. The steps for classifying additional images through the CODAvision GUI are detailed in Procedure 5. Users can select the images to segment, choose a pretrained model for inference and optionally modify the color palette for the colorized segmentation masks or perform additional object-based morphology analysis.
Description of the expertise needed to implement the protocol
This protocol is designed for researchers with some training and experience in computational biology. In particular, users must possess some knowledge of anatomy and coding to use this workflow, which we describe here.
Successful implementation of this workflow requires knowledge of anatomical structures and how they appear in histology/radiology images. For example, a user wishing to segment cancer metastases in H&E images of mouse lung must understand how to differentiate cancer cells from the functional cells of the lung in these images. The user will use this knowledge to manually annotate and qualitatively assess model performance. We provide a ‘Supplementary annotation guide’ (Supplementary Information) that gives some background on how to identify structures in medical images, which may be beneficial to some users.
Users must also possess a basic understanding of programming. While operation of this workflow is streamlined with a user-friendly GUI, initial installation requires familiarity with Python scripting, CUDA and cuDNN setup for GPU acceleration. Detailed instructions for these steps are provided on the GitHub page, where all codebase and dependencies are hosted. We suggest that users without programming knowledge obtain assistance when initially installing the package, after which operation of the GUI can proceed without significant coding expertise.
Limitations
This workflow, while powerful, possesses several limitations, which we document here. First, some of the model architectures included (DeeplabV3+ and U-Net) may require significant computational resources, such as high-performance GPUs. CODAvision may be used on computers without GPUs or on standard laptops, although we note that the speed of the workflow is impacted, as presented in Table 1.
Table 1 |.
Runtime comparison across different systems for image processing, training, inference and quantification steps using the demo dataset available
| Dell PC | Mac Studio | Lenovo LOQ 15ARP9 | MacBook Pro | |
|---|---|---|---|---|
| Windows 11, Intel Xeon W7-3455 CPU, 256 GB RAM, NVIDIA RTX 5000 Ada Generation GPU | macOS Sequoia 15.3, Apple M4 Max chip, 128 GB RAM | Windows 11, AMD Ryzen 7 7435HS CPU, 24 GB RAM, NVIDIA GeForce RTX 4070 Laptop GPU | macOS Tahoe, Apple M4 Pro chip, 24 GB RAM | |
| Downsampling images | 00:01:44 | 00:00:31 | 00:01:42 | 00:01:07 |
| Loading annotations | 00:01:35 | 00:00:46 | 00:01:36 | 00:00:55 |
| Creating tiles | 02:08:02 | 01:04:48 | 01:54:11 | 01:17:15 |
| Training model | 00:22:29 | 01:42:58 | 00:58:20 | 03:29:56 |
| Testing model | 00:00:39 | 00:00:24 | 00:00:50 | 00:00:43 |
| Classifying images | 00:01:40 | 00:00:43 | 00:04:00 | 00:02:36 |
| Quantifying images | 00:00:02 | 00:00:01 | 00:00:03 | 00:00:01 |
| Total time (h:min:s) | 02:36:11 | 02:50:11 | 03:00:42 | 04:52:33 |
Second, the segmentation models described here require highly specific manual annotations for training data. The benefit of this approach is the ability to, in the span of a few days, train highly accurate and highly customizable segmentation models in any cohort. The limitation is that a model trained on one organ (for example, mouse lungs) is not easily adapted to another organ (for example, human pancreas). Instead, users must generate new manual annotations for each new application.
Materials
Equipment
A computer with at least 16 GB of random access memory (RAM)
An NVIDIA GPU with at least 8 GB RAM or more (recommended)
An up-to-date operating system (Windows 10/11, OSX 11)
At least 2.5 GB of storage space
A working CUDA and cuDNN installation (instructions available at the provided GitHub page)
In the analysis described here we used a computer with the following specifications:
Workstation with 256 GB RAM and an NVIDIA GeForce RTX 5000 Ada Generation GPU running on Windows 11
Software
CODAvision software available in the following repository: https://github.com/Kiemen-Lab/CODAvision
Python Interpreter (for example, PyCharm, Visual Studio or Spyder)
- Image Annotation Tool (choose one of the two compatible software):
- Aperio ImageScope: a Windows-based viewer developed by Leica Biosystems, primarily used for viewing and annotating whole-slide images (WSIs). This software can be installed from the following link: https://www.leicabiosystems.com/digital-pathology/manage/aperio-imagescope
- QuPath: an open-source, cross-platform application for digital pathology and WSI analysis. This software can be installed from the following link: https://qupath.github.io/
Example dataset: Dataset 1: mouse lung histology, available at the following data repository link: https://doi.org/10.7281/T1W3JEFA
Procedure 1
▲ CRITICAL This protocol assumes that users possess a dataset of images intended for semantic segmentation and quantification. For sample datasets, see Dataset 1 in the ‘Software’ section of this protocol and see the analysis for it in Extended Data Fig. 4. Furthermore,we demonstrate that CODAvision can be applied to diverse image types, including histology, MRI and computed tomography. Our primary demonstration, and the sample datasets provided, are histological images scanned at 20× magnification (~0.5 μm per pixel resolution), though the procedure could be similarly applied to images scanned at higher or lower resolution.
Constructing a training dataset for deep learning training
● TIMING 8–10 h
▲ CRITICAL This procedure describes the annotation protocol in detail for Aperio ImageScope. Users wishing to annotate using QuPath may do so following analogous steps and noting differences highlighted in critical notes below.
-
1
Select six images to annotate from the initial cohort.
▲ CRITICAL The images selected for annotation should reflect the heterogeneity of the larger dataset. This may include selecting images from different scientific groups (control versus experimental conditions), images possessing distinct anatomical features and images with technical heterogeneity such as variation in lighting or focus. Construction of a heterogeneous training dataset will improve the robustness of the segmentation.
-
2
Create a folder named ‘Training dataset’ and another folder named ‘Testing dataset’.
-
3
Copy five of the selected images to the ‘Training dataset’ folder and save the sixth image in the ‘Testing dataset’ folder.
-
4
Install Aperio ImageScope following the installation instructions available at: https://www.leicabiosystems.com/digital-pathology/manage/aperio-imagescope.
-
5We suggest users change two settings in Aperio Imagescope upon installation to improve the user experience.
- To increase the maximum allowable zoom for precise annotation, navigate to Tools > Options > General Tab > Maximum magnification and enter 1000%.
- To automatically save annotations when exiting the program, navigate to Tools > Options > Annotation Tab > Annotation Settings, and check the box ‘Automatically save annotation changes’.
-
6
Open one of the images from the ‘Training dataset’ folder in Aperio ImageScope.
-
7
Create the annotation layers by navigating to View, then Annotations to show the ‘Annotations - Detailed View’ window and click the ‘+’ button to add an annotation layer. Rename the annotation layer by clicking on the layer name’s top and press ‘F2’, then input the desired name.
-
8
Create one annotation layer for each object you would like to train a model to segment. Once all layers are created, press save. This will generate an .xml file corresponding to this image that will contain the annotation coordinates that will be used for model training.
▲ CRITICAL All annotated images must have the same layer order. Create a layer for every structure, even in images where those structures are absent.
▲ CRITICAL The testing dataset must contain at least one annotation of each annotation layer and at least 15,000 annotated pixels per class to get a good assessment of the model’s accuracy (pixel count will be outputted automatically by CODAvision during model testing, as described in Procedure 3). If all layers are not present in any single image, consider using multiple images for testing so that the overall testing dataset contains at least one annotation of each annotation layer.
▲ CRITICAL The .xml files in the training and testing folders must correspond to the images they annotate, with identical filenames. If an .xml file has a different name than its associated image, the code will fail. Ensure that every image has a matching .xml file and vice versa.
▲ CRITICAL If QuPath is used for annotation, ensure that both the nesting structure and layer order are consistent across all images. To integrate these annotations into the workflow, export each image’s annotations as a GeoJSON file via File → Export objects as GeoJSON. Next, convert the GeoJSON files to XML format using the provided conversion script available at the following GitHub repository, which includes detailed usage instructions: https://github.com/Kiemen-Lab/GeoJSON2XML. After generating the XML files, ensure they are placed alongside their corresponding images, following the directory structure required by the standard protocol and continue the workflow as instructed.
Choosing which structures to annotate
▲ CRITICAL The number of annotation layers you should generate depends on your research objective. In histology, many cell types can be differentiated by the trained eye, including various epithelial, vascular and stromal compartments. We suggest that users first make a list of the major structures present in their images, then group these structures until the desired granularity is obtained. See Fig. 3 for an example of a high-detail and a low-detail model trained on fetal rhesus macaque kidney histology. Where exhaustive anatomical labeling is desired, the user can generate a highly specific list of structures identifiable in H&E (Fig. 3a). For a more focused project, the user can group labels to reduce the number of annotation layers and increase the speed of the project (Fig. 3b).
Fig. 3 |. A sample histological image of fetal rhesus macaque kidney with anatomical annotations overlaid.

a, For a high-detail model, 17 tissue structures are identifiable in the kidney. b, For a vasculature-focused model, the annotation layers can be grouped to remove unnecessary labels23.
▲ CRITICAL No matter your research question, all models should contain a background or whitespace layer in the annotation dataset to contain nontissue pixels in the image.
-
9
Begin annotating using the ‘Pen’ tool. Start the annotation by clicking and annotating the area of interest in the main window. Use the pen tool to manually outline the region of interest, closing the drawn shape once complete. If you wish to modify the annotation, click and redraw over the annotated region until you achieve the desired level of precision in refining the borders.
▲ CRITICAL We recommend that users periodically save the annotations manually by clicking the save button in the ‘Annotations - Detailed View’ window. This will help prevent potential data loss.
-
10
For a high likelihood of generating a high-accuracy first segmentation model, make ~20 annotations per tissue structure per image (training and testing), noting that, for rare structures, there may be fewer than 20 examples per image. To obtain the first model results more rapidly, the user may make ~10 annotations per tissue structure per image, acknowledging that the model may require more iterative finetuning to achieve high accuracy results.
▲ CRITICAL We provide guidance on good annotation practices. For more detailed notes on annotating, see the companion annotation guide in the Supplementary Information.
-
11
For nesting, make overlapping annotations of different classes following a consistent nesting hierarchy (Fig. 4a).
▲ CRITICAL The resulting quality of your segmentation model relies on the quality of your annotations. Zoom in to high magnification when annotating and aim to annotate the structure boundaries very cleanly and consistently (Fig. 4b).
▲ CRITICAL Define the nesting hierarchy before beginning annotations. This hierarchy must remain consistent across all annotated images and will be used during the deep learning model parameterization described in Procedure 2.
Fig. 4 |. Annotation hierarchy and quality control for CODA deep learning model training.

a, A hierarchical nesting diagram illustrating tissue classification levels. b, A comparison of adipocyte annotation precision within pancreatic acinar tissue: suboptimal annotation (left) versus optimal annotation (right). c, An iterative annotation workflow demonstrating the targeted addition of annotations in misclassified regions to enhance model performance. PanIN, pancreatic intraepithelial neoplasia.
Nesting
▲ CRITICAL We have observed that CODA segmentation models yield better results when the tissues are identified within their microenvironments. To achieve this, employ an annotation technique called ‘nesting’. Nesting uses a hierarchical tissue organization in which higher-level tissues can be ‘nested’ inside lower-level ones. Figure 4a shows the arrangement of three hypothetical annotated types: tissue A (triangles), tissue B (squares) and tissue C (circles). Varying the nesting hierarchy changes how these annotation layers are imported for model training.
-
12
Annotate structures across the entire image, not just in one region.
-
13
Include diverse examples of all annotation classes. Building a dataset with varied morphologies for each class ensures optimal performance of your model on unseen data.
-
14
Include ‘non-ideal’ annotations for each class to enable the model to correctly classify tissue types even in the presence of noise. For example, annotate structures that are slightly blurry, darker, or paler.
-
15
When glandular structures that contain a lumen are present, include background annotations within the lumen if noise, such as fluid or red blood cells, is present (Fig. 4a, right).
-
16
Include five to ten annotations at tissue structure edges when annotating whole slide images to ensure accurate differentiation between tissue borders and background during classification.
-
17
When re-annotating to improve model performance, first review the classified images to identify regions of misclassification. Focus your annotations on these regions to efficiently correct the model (Fig. 4c).
▲ CRITICAL Refer to the ‛Supplementary annotation guide’ (Supplementary Information) for detailed examples of tissue annotations and best practice.
Procedure 2: defining model parameters using the GUI
● TIMING 5 min
-
1
After completing the annotation step, download CODAvision by following the README instructions present on the following GitHub page: https://github.com/Kiemen-Lab/CODAvision.
◆ TROUBLESHOOTING
-
2
If desired, import the Python code to an integrated development environment.
-
3
Run the CODAvision.py code to execute the GUI to parametrize the settings for the model.
◆ TROUBLESHOOTING
▲ CRITICAL In the GitHub repository, there is also a command-line script (CODAvision/base/scripts/non-gui_workflow.py) with detailed documentation for dataset parameterization without using the GUI. This script has been successfully tested on the SciServer cloud infrastructure at Johns Hopkins University, demonstrating the feasibility of cloud-based execution of the workflow.
File location tab (Fig. 5a)
Fig. 5 |. CODAvision GUI tabs.

a, The file location tab: configuration of dataset paths, model name and training image resolution. b, The segmentation settings tab: parameterization of whitespace management c, The nesting tab: configuration of nested annotation hierarchy for overlapping tissue annotations. d, The advanced settings tab: modification of CNN hyperparameters and selection of annotation classes for component analysis.
-
4
Browse for the folder containing the training annotations. This folder should contain the annotated images and .xml files generated during Procedure 1.
-
5
Repeat Step 4 to browse for the folder containing the testing annotations.
-
6
Enter a desired name for the deep learning model. By default, the name is prepopulated with today’s date but may be customized.
-
7
Specify a desired image resolution by selecting an option from the dropdown list.
Choosing a training resolution
▲ CRITICAL The choice of training resolution is critical for achieving the desired segmentation accuracy while managing computational resources efficiently. For cellular-level analyses, we recommend using 10× (1 μm per pixel), whereas organ-level or large tissue structures can be effectively segmented at 1× (8 μm per pixel). The key consideration is the trade-off between segmentation detail and computational efficiency. Higher resolutions provide finer detail but result in larger file sizes and extended processing time, and require higher precision manual annotations. CODAvision also enables users to provide pregenerated downsampled images and to input this custom scale factor instead of choosing from one of the predefined resolutions. To do this, select ‘custom scale’ and browse for the folder containing the scaled .tif or .png images.
-
8(Optional) Select ‘Custom’ from the ‘Resolution’ dropdown menu to train on resolutions different from the default options. This action will display a scaling factor input field where you can specify the desired downsampling ratio (must be ≥1) (Fig. 6).
- To use different images for downsampling instead of the annotated images, select the ‘Scale custom images’ checkbox. Then, locate the custom image directories, which should be organized into separate training and testing folders. The image filenames must correspond exactly to their respective annotation files
-
If your custom images are prescaled, activate the ‘Use pre-scaled images’ checkbox. Ensure that these images match the value specified in the ‘Scaling factor’ field and are .tif or .png filetype▲ CRITICAL Choose this custom option described in Step 8 when working with images not originally annotated in the recommended formats (.ndpi or .svs). The custom image downsampling feature accepts the following file formats: .ndpi, .svs, .tif, .jpg, .png and .dcm.
-
9
After completing all sections on the ‘file location’ tab, click ‘Save & Continue’ and move to tab 2: ‘segmentation settings’ (Fig. 5b). A popup window will appear, prompting you to either automatically import annotation data from a random XML file or manually select a specific XML file to import the data from.
-
10
Once an XML file is selected, the interface will advance to tab 2: ‘segmentation settings’. This tab serves as the interface for defining pixel background management for each annotated class and will be prepopulated with the annotation classes and colors defined during manual annotation.
Fig. 6 |. Optional tab for custom downsampling of images or for providing pre-downsampled files to the GUI.

Selection of the ‘Custom’ option from the resolution dropdown menu is shown on the left. The checkbox used to scale custom images relative to those used for manual annotation, based on the value entered in the ‘Scale factor’ text box, is shown on the right. Alternatively, users may select prescaled images to directly load images that have already been downsampled by the specified scale factor relative to the annotation images.
Segmentation settings tab
-
11
Determine how best to handle the white/background pixels (referred to as ‘whitespace’ in this protocol) for each annotation layer. To do so, click on an annotation class from the table, and select one of the three available options in the ‘Annotation class whitespace settings’ (see Fig. 7 for examples).
Fig. 7 |. Example annotations where each whitespace management option is best.

H&E histology structures for which removal of background pixels is recommended, including blood vessels, bile ducts and stromal regions, are shown on the left. Structures for which background pixels should be retained as part of the segmented region, such as adipocytes, are shown in the middle. Examples for which we recommend retaining both background pixels and tissue regions, as pixels with background-like appearance may correspond to structures of interest, including nerves, hepatocytes and pancreatic islets, are shown on the right.
Managing background in manual annotations
▲ CRITICAL Proper whitespace management is critical for training high-accuracy deep learning models. Here, we provide advice on when to automatically remove whitespace or non-whitespace from your annotation layers:
-
12
Select ‘Remove whitespace’ to eliminate the background pixels from an annotation layer. This is relevant for scenarios such as excluding the lumen from a glandular structure or to remove the white pixels intermixed between stromal fibers. Select ‘Keep only whitespace’ to retain only the background pixels in the annotation. This is relevant when annotating fat and aiming to exclude non-white lines separating individual fat cells. Select ‘Keep tissue and whitespace’ to retain both background and non-white pixels. This is appropriate for the noise/background layer, as the annotated regions may contain both whitespace and noise such as shadows, or when annotating a solid structure such as hepatocytes in the liver. Refer to Fig. 7 for visual examples.
-
13
Upon selecting an ‘Annotation class whitespace settings’ option, click ‘Apply’ or ‘Apply all’ to update the table.
-
14
Define the destination class for whitespace pixels removed from annotation layers where the option ‘Remove whitespace’ was selected. In general, the destination class should be the background class.
-
15
Similarly, define the destination class for removed non-whitespace pixels taken from the annotation layers where the option ‘Keep only whitespace’ was selected. In general, the destination class should be the stromal class. This input must be defined even if no annotation layer was assigned with ‘Keep only whitespace.’
-
16
(Optional) To change the color assigned to any annotation layer, select the desired annotation class from the table and click the ‘Change Color’ button. In the color picker window, select the desired color, then click ‘OK’ to confirm the color change.
▲ CRITICAL By default, the model’s classification output will use the same colors as those used during manual annotation (shown as background colors in the table). Color changes are purely aesthetic and do not affect model performance, but well-chosen colors may improve users’ ability to visually interpret and present the segmentation results. To ensure accessibility, we recommend selecting color palettes that are friendly for color-blind users.
-
17
(Optional) To combine annotation layers, hold the ‘Ctrl’ key and select the desired rows in the table. Click the ‘Combine classes’ button and, when prompted, enter a name for the combined class. In the color picker window, select a color for the combined class.
-
18
(Optional) To delete unwanted annotation layers, select the appropriate row by clicking inside the table and then click the ‘Delete class’ button.
▲ CRITICAL Delete unwanted annotation layers that should not be included in the model training. This is useful for removing empty layers.
-
19
(Optional) Click the ‘Reset list’ button to return the annotation class table to its default state. This will remove whitespace management choices, uncombine any combined layers and restore deleted layers.
-
20
Once this tab is completed, click ‘Save & Continue’ in the bottom right corner. The interface will automatically advance to tab 3: ‘nesting’.
Nesting tab
-
21
Define the appropriate nesting order to ensure correct handling of overlapping annotation classes (Fig. 5c). This order should have been established during the manual annotation step in Procedure 1. To define the nesting order in the GUI, select a desired annotation layer. Use the ‘Move Up’ and ‘Move Down’ buttons to adjust its position according to layering priority.
▲ CRITICAL The ‘nesting’ tab allows the user to configure the layering hierarchy of annotation classes from lowest to highest priority in cases of overlap. The class at the top of the table is assigned the highest priority, while the class at the bottom is assigned the lowest priority.
-
22
(Optional) To define different nesting orders to several annotation layers that were combined in the segmentation settings tab, check the ‘Nest uncombined data’ box.
▲ CRITICAL If annotation classes were combined in the previous tab, the nesting table will display these combined layers by default. If a combined layer is made from two layers that require different nesting priority, check the ‘Nest uncombined data’ box. Refer to the annotation guide in the Supplementary Information for detailed examples on establishing the nesting order.
-
23Once the nesting tab is completed, the user has four options:
- (Most likely) Select ‘Save and train’ to immediately proceed with model training
- Click ‘Save and close’ to save the model configuration but NOT train the model.
- Click ‘Continue to advanced settings’ to define more complex model parameters before training the model. This tab will enable the user to adjust model hyperparameters, select the model architecture, or select classes for advanced quantitative analysis.
- Select ‘Return’ to go back to tab 2.
Advanced settings tab
-
24
(Optional) To adjust the default tile size input to the segmentation model, choose a new tile size from the dropdown list (default: 1,024 × 1,024 RGB) (Fig. 5d).
▲ CRITICAL The training tile size default is 1,024 × 1,024 pixels. Users with GPU, central processing unit (CPU) or RAM constraints should consider a smaller size such as 512 × 512 or 256 × 256 (the size must be a power of 2). If the computer does not have a GPU, the training pipeline will automatically fall back to CPU training; in this case, tile size should be adjusted accordingly to available RAM.
-
25
(Optional) To adjust the training and validation tile number, click on the up and down arrows next to each respective text box. The default number of training tiles is 15, and the default number of validation tiles is 3. Users may increase these numbers for especially large (>50 annotated images) or small (<5 annotated images) training datasets.
-
26
(Optional) Adjust the number of images used for tissue mask thresholding (described in greater detail in Procedure 3) by clicking on the up and down arrows next to the ‘Tissue mask evaluation:’ text box. By default, this number is three images, but a higher number may be desired for larger or more diverse datasets.
-
27
(Optional) Click the ‘Individual tissue mask evaluation’ checkbox to customize the tissue mask threshold for each image in the dataset. This option is recommended when image appearance varies substantially across the dataset and images may have varying background intensity thresholds.
-
28
(Optional) Click the ‘Remake tissue mask’ checkbox to regenerate the tissue masks during retraining. This ensures that mask creation is not skipped in subsequent runs.
-
29
(Optional) Choose the desired model architecture from one of two available options in the dropdown box (DeepLabV3+ or UNet). By default, the training architecture is DeeplabV3+.
Choosing a model architecture
▲ CRITICAL The choice of model architecture depends on computational resources available and the specific structures to be segmented. UNet contains ~41 million parameters, and in our benchmark test, required ~40 min to train on an NVIDIA GeForce RTX 5000 Ada Generation GPU. By contrast, DeepLabV3+ contains ~12 million parameters, and trained in ~75% of the training time of UNet. While UNet excels in capturing fine-grained details owing to its deeper architecture, DeepLabV3+ is often more computationally efficient and may be preferable for users with limited computational resources or time constraints. Users are encouraged to experiment with both architectures to determine which best suits their specific needs.
For computationally experienced users, additional architectures can be integrated into the workflow by modifying the codavision/models/backbones.py code, where the provided networks are implemented as Python classes. Users can also adapt the codavision/models/training.py and codavision/CODA.py to enable training on the new architecture and include it as an option in the dropdown menu of the advanced settings tab in the GUI.
Adding additional model architectures to CODAvision
▲ CRITICAL CODAvision follows a plugin-based architecture that enables seamless integration of new segmentation backbones. To add a model, first implement the desired architecture as a new Python class in base/models/backbones.py, following the structure of the existing BaseSegmentationModel abstract. The new class must define a build_model() function that constructs and returns the model. Second, register the new class in the model_call() function of the same file by adding a new case with the chosen model’s name. Third, ensure that base/models/training.py can recognize the new backbone: the SegmentationModelTrainer loads the model via model_call(), assigns the loss function and compiles with the chosen optimizer under the current distribution strategy. Finally, add the model name to the dropdown options in gui/components/ui_definitions.py, which queries get_available_models() to expose all registered models in the advanced settings menu, so the architecture is selectable through the GUI. Additional detailed README instructions on how to add additional model architectures can be found in the GitHub repository.
-
30
(Optional) Modify the batch size used for training. By default, the batch size is set to 3.
▲ CRITICAL Adjust the batch size according to your memory capacity. Larger batch sizes accelerate training but require more memory, while smaller batch sizes increase training duration but reduce memory requirements. Monitor memory usage (GPU, CPU and RAM) during initial training attempts to optimize this parameter for your system.
-
31
(Optional) Select classes for detailed quantitative analysis (object number count and size per image).
▲ CRITICAL For a selected class in each segmented image, the total number of objects >500 pixels will be counted and the object size in pixels will be documented. These data will be exported in a .csv file following model training and image segmentation.
-
32Once the advanced settings tab is completed, the user has three options:
- (Most likely) Select ‘Save and train’ to proceed with model training.
- Select ‘Save and close’ to save the configuration but NOT train the model.
- Select ‘Return’ to go back to tab 3.
Procedure 3: image preprocessing and model training
Tissue mask thresholding
● TIMING 5–10 min
-
1
Upon completing the parameter configuration in the CODAvision GUI, the software will initiate execution and begin downsampling the training images to the resolution specified in Procedure 2, Step 7.
-
2
Next, a popup window will appear, enabling interactive selection of a threshold value to separate tissue pixels from background pixels (Fig. 8).
-
3
In the popup window, an image will be displayed. Double click on a region of the image containing both tissue and whitespace.
-
4
The popup window will reload and display a magnified view centered on the selected region. Confirm the selection or choose a different region until a suitable area with sufficient tissue and background is identified.
-
5
Once the region is confirmed, a new window will prompt the user to adjust the threshold cutoff until the background is detected optimally.
-
6
After selecting the desired threshold, another full-size image will load to repeat the process until the desired number of images has been assessed. The average threshold value from this process will then be applied to all images in the cohort.
▲ CRITICAL Tissue mask threshold optimization is required only once. If retraining the model on the same dataset, a popup window will prompt the user to either ‘Keep current tissue mask evaluation’, which reuses the previously determined threshold and skips this step, or ‘Evaluate tissue mask again’, which initiates a new round of threshold selection. If the ‘Individual tissue mask evaluation’ checkbox (Procedure 2, Step 27) was selected and the user chooses to ‘Evaluate tissue mask again’, a prompt will appear allowing manual selection of specific tissue masks that the user desires to remake. These tissue masks are located in the training annotation path in the downsampled images folder. Alternatively, selecting ‘Redo all images’ will apply the thresholding process to the entire dataset.
Fig. 8 |. The tissue threshold selection interface.

a, Region of interest selection in a lung WSI to optimize threshold values. b, Threshold adjustment interface with optimal tissue–background discrimination demonstrated (bottom).
CODAvision image processing and model training
● TIMING 2–3 h
-
7
Following successful completion of tissue mask optimization, the image preprocessing and model training will proceed automatically (Fig. 9).
▲ CRITICAL This step of the protocol is completed automatically by the GUI once the training is initiated in Procedure 2. The user has no direct steps in this section, but the steps followed by the code are described below for clarity.
▲ CRITICAL Depending on the integrated development environment used to run the CODAvision script, users may need to close any popup plots, such as the confusion matrix generated in Step 13 or the segmented image displayed in Step 14, for the code to finish execution.
◆ TROUBLESHOOTING
-
8While executing, the code will output several text statements to the command window, in the following order:
- The specified training and testing paths
- The model name and selected resolution
- The list of annotated classes and their associated colormap values
- A confirmation message indicating that the model metadata and classification colormap have been saved
- For each annotated image, a message will appear indicating the status of the annotation metadata import:
- If the annotation data were loaded in a previous training and the annotations and model parameters have not changed, the message will state that the metadata for this image was previously loaded
-
If the annotation data are being loaded for the first time, the following checkpoint messages will be displayed: (1) confirmation that the annotation data have been imported from the .xml file, (2) notification that an annotation mask has been generated and saved as a .jpg file and (3) confirmation that smaller image patches have been cropped around each annotated region using the annotations mask, producing ‘bounding boxes’ of the RGB image and mask for downstream processing◆ TROUBLESHOOTING
-
9
After loading all annotation metadata, the code will output the total size of the training dataset, including the percentage of the training dataset contributed by each class.
▲ CRITICAL It is recommended to have a well-balanced dataset with many examples of each class. The codes will automatically calculate the percentage of annotations for each class. We suggest that the minimum class should make up at least 5% of the annotated pixels for the most annotated class to ensure sufficient heterogeneity is provided for each annotation label and that the heterogeneity of the annotation classes is represented roughly equally. If any class is under the 5% threshold, we recommend the user attempt to add manual annotations of the least prevalent class through further annotation of the training images or through addition of new training images.
-
10
After the metadata is loaded for all annotation images and the composition of the training dataset is calculated, training and validation tiles will be constructed.
Fig. 9 |. The visualized workflow of Procedure 3.

The diagram shows the workflow for Procedure 3.
How to generate robust deep learning models
▲ CRITICAL Construction of a large and diverse training dataset is important to generate robust deep learning models. We augment the manual annotations in several ways to increase the heterogeneity of the training data. For each training tile, a large zero-value image of size 10,000 × 10,000 × 3 pixels is generated. Bounding boxes containing processed manual annotations are randomly added to the tile until the tile is >55% filled. For each iteration, the least represented class in the tile is determined, and an annotation bounding box containing that class is added to the tile. Approximately 50% of tiles are augmented by adjusting the hue (scaling RGB channels independently within a range of 0.88–1.12), scaling (within a range of 0.6–0.95 for downscaling and 1.1–1.4 for upscaling), rotating (at random angles from 0° to 355° in increments of 5°) and applying a Gaussian blur (with sigma values randomly selected from a predefined set, for example, 1.05, 1.1, 1.15 and 1.2). Each bounding box is thus added many times to the training tiles, but each time in a new position surrounded by different neighbors, and augmented in several ways. Once >55% of the tiles are filled, these large tiles will be cropped to a size of 1,024 × 1,024 × 3 pixels (or the custom size if defined by the user in the advanced settings tab of the GUI) and saved.
-
11
After the predefined number of training and validation images have been created, model training will start. All training and validation tiles will be normalized through zero-centering. The model is validated three times per epoch, and training stops on the basis of an early stopping patience parameter specified in the ‘CODAvision/base/config.py file’.
◆ TROUBLESHOOTING
▲ CRITICAL If the computer running CODAvision does not have a GPU, the CNN training will automatically fall back to the CPU, increasing the runtime accordingly.
▲ CRITICAL To account for any class over-representation, the CNN training pipeline has incorporated a weighted sparse categorical cross-entropy loss function to mitigate any bias.
-
12
Following model training, annotation masks will be generated for the test dataset as detailed in Step 8 of this Procedure.
◆ TROUBLESHOOTING
-
13
The testing image(s) will be segmented, with the resulting segmentation masks and colorized overlays saved as .tif and .jpg files, respectively. This output will be compared with the metadata collected in Step 12 of this Procedure to generate a confusion matrix.
▲ CRITICAL A minimum of 15,000 pixels per annotated class is recommended for the ground truth annotations used in the testing dataset. Pixel counts will be displayed when computing the confusion matrix after training.
-
14
The images specified in the training images folder will be segmented, with the resulting segmentation masks and colorized overlays saved as .tif and .jpg files, respectively.
-
15
The bulk tissue composition of the images in the training images folder will be calculated and saved in a .csv file.
-
16
(Optional) Component analysis will be calculated for selected annotation classes in the ‘advanced settings’ tab described in Procedure 2 and saved in a .csv file. This step would be skipped if no classes were selected.
Generated outputs for Procedure 3
Model hyperparameters, trained weights and results will be stored in the training path file location. A summary of input and output file formats for the CODAvision workflow is provided in Fig. 10. This workflow generates outputs per annotated image, per trained model and per segmented image.
Fig. 10 |. An overview of CODAvision input and output data formats.

The input files accepted by CODAvision are shown on the left. The output files generated for each segmented image, model and dataset are shown on the right.
-
17‘data_py’ folder:
- Subfolder for each training image containing:
- ‘annotation.pkl’: processed vertex coordinates corresponding to the manual annotation data saved in each .xml file.
- ‘view_annotations.png’: grayscale mask images containing pixel labels for each annotated region, considering the ‘whitespace’ management and the nesting order.
- ‘model_name_boundbox’ subfolder: contains cropped RGB and mask images corresponding to each annotated region identified in ‘view_annotations.png’.
- ‘im’ subfolder: contains the RGB copy of the bounding boxes.
-
‘label’ subfolder: contains the corresponding mask of the bounding boxes.▲ CRITICAL The ‘data_py’ folder contains metadata used internally for image processing and training tile generation, not for analysis.
-
18
The ‘check_annotations’ folder: contains colorized masks of the training images with annotation patches highlighted in colors defined in tab 2 of the GUI.
-
19‘Model_name’ folder.
- ‘training’ folder: contains image tiles used for model training
- ‘validation’ folder: contains image tiles used for model validation
- ‘best_model_X.keras’: weights of the CNN for the best training epoch, where X is the name of the model architecture that was trained
- ‘X.keras’: weights of the trained CNN, where X is the name of the model architecture that was trained
- ‘net.pkl’: annotation parameterization settings
- ‘model_color_legend.jpg’: color map used in colorized segmented images
- ‘model_evaluation_report.pdf’: detailed performance report (Fig. 11)
-
‘confusion_matrix_X.png’: confusion matrix with precision, recall and accuracy metrics, where X is the name of the model architecture that was trained◆ TROUBLESHOOTING▲ CRITICAL Review the ‘check_annotation’ images to verify that your annotations are correctly overlaid on the downsampled images. Pay close attention to how the whitespace is handled (adjust the whitespace settings in tab 2 of the GUI if needed, or regenerate the tissue mask images). In addition, ensure that the nesting order is correct; if necessary, revise the annotations or modify the nesting settings in tab 3 of the GUI.▲ CRITICAL We suggest a minimum overall accuracy score >90%, and a minimum per-class precision and recall exceeding 85% for acceptable training results. For models that do not meet these standards, we suggest users add additional training annotations and retrain to improve the metrics.
-
20The downsampled image folder is named on the basis of the chosen resolution (for example, ‘1×’, ‘5×’, ‘10×’ or a user-defined custom scale).
- ‘TA’ subfolder: ‘tissue area’ masks for the training images.
- ‘classification_model_name’ subfolder, where ‘model_name’ is the user-defined name of the trained model. This subfolder will contain:
-
Grayscale segmentation masks of each training image in .tif format ‘check_classification’ subfolder and colorized segmentation masks in .jpg format.◆ TROUBLESHOOTING
- ‘image_quantifications.csv’: pixel count and tissue composition for each training image.
-
‘annotation_class_name_count_analysis.csv’: object count and size analysis for each class.▲ CRITICAL To visually assess the model accuracy, review the ‘check_classification’ images and search for areas that are misclassified.
-
Fig. 11 |. An example of a report.

A sample PDF report summarizing the performance metrics and training parameters of a model trained to recognize normal hepatic cells and cancer in human liver histology is shown.
Procedure 4: model optimization
● TIMING 1–10 h
▲ CRITICAL If a trained model or classified images are of unsatisfactory quality (overall accuracy <90%, minimum precision or recall <85%, or visually poor performance), follow this procedure to retrain the model efficiently. It is common when starting a new project to train two to three models in quick succession before achieving satisfactory results. Here, we describe four techniques to improve model performance (Fig. 12).
Fig. 12 |. Suggested solutions to incrementally improve model performance.

To optimize a pretrained model, first review segmentation overlays in the ‘check_classification’ subfolder to identify misclassified areas and annotation masks in the ‘check_annotations’ folder to identify potential errors introduced during manual creation of the training dataset. Next, update the training annotations by correcting previous labels or adding new annotations in regions of misclassification. Finally, adjust relevant training parameters through the GUI if needed, retrain the model and repeat the workflow iteratively until you achieve the desired accuracy and segmentation quality.
- First, determine whether there is human error in the GUI settings. If yes, reconfigure the settings and retrain the model without adding new annotations.
- Review the color overlay images inside the ‘check_annotation’ folder to determine whether the annotations contain errors or model parameters were defined incorrectly.
- Does the whitespace management look correct? For example, in an annotated blood vessel, are the background pixels inside the lumen correctly extracted as whitespace? If incorrect, consider adjusting the whitespace management in tab 2 of the GUI (see Fig. 7 for examples).
- Does the nesting order in the images look correct? For example, if noise was annotated inside of a duct, is that region defined in the check_annotation file as noise or as duct? If incorrect, consider adjusting the nesting order in tab 3 of the GUI (see Fig. 4 for examples).
-
If the nesting order and whitespace management look correct, is anything else noticeable in the file? Are any structures incorrectly annotated? Consult the ‘Supplementary annotation guide’ (Supplementary Information) for detailed best practices in annotation and correct these errors by editing the annotations where necessary.▲ CRITICAL Biological images contain complex and subtle structures. Human errors in manually annotating images are common, especially for new users. If errors are identified in the annotated files, correct them and retrain, and remember to also correct errors in the testing dataset.
- Review the tissue area masks to determine whether the threshold is correctly separating the tissue and background (see Fig. 8 for examples). If this threshold looks incorrect, delete the subfolder named ‘TA’ so that a new threshold may be calculated upon retraining.
- Second, view the images saved inside the folder named ‘check_classification’ to visually assess the model performance and identify misclassified regions. These files show a color overlay of the model segmentation across whole images.
- Review several images and identify common ‘patterns’ in misclassification. For example, stroma consistently misclassified as vasculature, or background pixels consistently misclassified as fat.
- Determine whether there are any structures present in the images that were missed during the previous model training and were not included as an annotation layer (for example, if blood vessels were not included as a label but are present in the images). If identified, determine whether these structures should be combined with a current annotation layer or added as an additional annotation layer.
-
After correcting any problems caused by human error in model parameterization and retraining the model, the user may improve the model results through further annotation. Focusing on the list of patterns in misclassification generated in Step 2 of this Procedure, add annotations to the identified problem areas.
▲ CRITICAL If the model frequently misclassifies unseen structures when applied to images beyond the training set, consider augmenting the training dataset with entirely new training images, rather than adding additional annotations to images previously used for training.
-
If all the human errors have been corrected and the training dataset is extensive, the user may also consider tweaking model parameters in the advanced settings tab of the GUI.
▲ CRITICAL When retraining a model using identical annotations (for example, same training and validation tiles) but with a different architecture, load previous model settings using the ‘Load prerecorded data’ button, set the model name to the one used to train the original model and modify the model architecture via the advanced settings tab in the CODAvision GUI (Procedure 2, Step 29). This approach eliminates the need to repeat Steps 1–10 of Procedure 3, allowing rapid retraining of models.
-
Repeat the steps of Procedures 1–4 to retrain the model and reoptimize the annotations and model parameterization until satisfactory results are obtained. When retraining a model, consider using the ‘Load prerecorded data’ option to prepopulate the model parameters.
▲ CRITICAL If data are loaded using the ‘Load prerecorded data’ button, the option to select a specific XML (Procedure 2, Step 9) file for importing annotation data will be bypassed. Instead, the settings from the previously saved model will be automatically applied to the ‘segmentation settings’ tab.
Procedure 5: image segmentation using a pretrained model
● TIMING 10–15 min
▲ CRITICAL After training a model, the user may segment additional images beyond the original training set. Models trained on manual annotations at a whole-organ label can be used for inference on the same anatomical regions and imaging modality. Models trained to distinguish structures within one organ should not be applied to segment other organs, nor should models trained on one imaging modality be used for another. See Extended Data Fig. 5 for an example of inference on an independent dataset (pancreatic WSIs obtained from the GTEx portal, accessed 30 October 2025).
▲ CRITICAL To ensure high model performance, only images with the same resolution as the training data set should be classified with a pretrained model.
Run the CODAvision.py code to restart the CODAvision GUI.
Load the training weights of the model you wish to use (Fig. 13a). To do this, either select ‘Load prerecorded data’ and select the folder containing the ‘net.pkl’ file you wish to use, or manually input text in the ‘Training annotations’, ‘Resolution’ and ‘Model name’ fields. When these fields are completed and a trained model is detected, a green ‘Classify images’ button will appear. Click this button to proceed to the ‘segment images’ tab.
After proceeding to the ‘segment images’ tab (Fig. 13b), input folders containing images you wish to classify. Click the ‘Browse’ button to add a desired folder to the table list and repeat until all desired folders are added.
-
(Optional) Modify the segmentation color map. Change the colors of annotation classes in the table on the left. This new color overlay may be visualized in real time to ensure readability through application to the sample image displayed in the center of the tab. The updated color map will be used to generate the colorized segmented images.
▲ CRITICAL If the objective is only to modify the color map of previously classified images, input the root directory path of images that have been classified by the model specified in Step 3, and click ‘Apply’. The code will begin execution, bypassing the segmentation, and will change the colorized segmented images.
(Optional) Perform annotation class component analysis. Select the annotation classes to quantify in the table on the right.
-
Click ‘Apply’ to segment all new images in all folders added to the segmentation table and to perform the quantitative analyses as defined.
▲ CRITICAL If the objective is only to modify the color map of previously classified images, input the root directory path.
▲ CRITICAL If the qualitative assessment of the inference result on an independent dataset has an unsatisfactory quality, consider following the model optimization steps explained in Procedure 4 to fine-tune the segmentation.
Fig. 13 |. The GUI tab for classifying additional images with pretrained models.

a, The file location tab showing the ‘Classify images’ button, which appears when a trained network is detected in the specified data folder. b, The interface for selecting new image set paths to classify with the pretrained model, modifying segmentation mask color map and performing additional annotation class component analysis.
Troubleshooting
Troubleshooting information can be found in Table 2.
Table 2 |.
Troubleshooting table
| Step | Problem | Possible reasons | Solution |
|---|---|---|---|
| Procedure 2 | |||
| Step 1 | Unable to install cudatoolkit and cuDNN dependencies | SSL certificate verification issues | Make sure that OpenSSL is installed and properly configured on your system. You can download and install OpenSSL from https://slproweb.com/products/Win32OpenSSL.html. After downloading OpenSSL add its bin directory to the environment variables |
| Step 1 | Unable to download CODAvision | Git not installed on computer | Download Git from the following website (https://git-scm.com/downloads/win) and restart your computer |
| Step 3 | Cannot execute CODAvision.py | Dependencies not properly installed | Ensure .dll path for OpenSlide is located in environment paths |
| Procedure 3 | |||
| Step 7 | Model training failure | Existing model with same name in file directory | Rename the model or delete the previous model’s folder |
| Step 8 | Code cannot identify annotation .xml file for training mask | Filename mismatch between annotated images, .xml files and downsampled .tif images | Ensure all filenames are consistent and rerun the code |
| Step 11 | GPU out of memory | Model architecture or batch size too resource-intensive | Choose a lighter model architecture or reduce batch size to match computer specifications |
| Step 11 | Validation loss increases after a certain number of epochs (overfitting) | Insufficient or incorrectly annotated training data | Increase number of annotated images or improve annotations |
| Step 12 | Model testing failure | Testing annotations lack classes present in training dataset | Include annotations for all classes in the training dataset |
| Step 19 | Confusion matrix shows all pixels classified into a single class | Numerical instability during GPU training | Retrain and test the model using CPU processing instead |
| Step 19 | Confusion matrix shows low precision and recall for light contrast classes similar to background | Poor tissue background cutoff | Repeat CODAvision workflow with different tissue background cutoff value |
| Step 20 | Misclassification in images outside training dataset | Insufficient annotation class diversity or heterogeneity | Include more annotation classes for morphologically different structures, create more heterogeneous annotations or increase training dataset with additional annotated images (refer to the ‘Supplementary annotation guide’ in the Supplementary Information) |
Timing
▲ CRITICAL The estimated processing times provided below are based on the training of the sample dataset provided (mouse lung histology), trained at the selected magnification of 5× (2 μm per pixel resolution), performed using a workstation equipped with an NVIDIA GeForce RTX 5000 Ada Generation GPU running Windows 11. These estimates reflect the processing time for an experienced CODAvision user using a GPU-equipped computer. Processing times will increase substantially if the software is run without GPU support. Table 1 presents a comparison of the runtimes for the entire CODAvision workflow on different laptops and PCs, including both GPU-accelerated training and CPU-only execution. PCs used a training tile size of 1,024 × 1,024 pixels while laptops used a training tile size of 512 × 512 pixels owing to memory constraints, which can be adjusted using the ‘advanced settings’ tab of the GUI (Procedure 2, Step 24). Note that users following the CODAvision workflow for the first time may require additional time to familiarize themselves with the best annotation practices and GUI guided parametrization.
Procedure 1
Steps 1–10, building annotation dataset: 8–10 h
Procedure 2
Steps 1–32, GUI-guided parameterization: 5 min
Procedure 3: image preprocessing and model training
Steps 1–6, tissue mask thresholding: 5–10 min
Steps 7–16, CODAvision image processing and model training: 2–3 h
Procedure 4: model optimization
Steps 1–5, model optimization: 1–10 h
Procedure 5
Steps 1–6, segmenting an image with a pretrained model: 10–15 min
Anticipated results
In this section, we provide use cases to demonstrate the value of the CODAvision outputs across several research areas. To demonstrate the versatility of CODAvision, we applied the workflow to four distinct research tasks (Fig. 14). For each application, we trained a model using DeepLabv3+ ResNet50 model architecture. Below, we briefly describe each use case and the analyses enabled.
Fig. 14 |. Example applications of CODAvision to different medical image types.

a, The quantification of metastatic burden in mouse lung and liver through semantic segmentation and automated tissue quantification. b, The deconvolution of Visium spots using CODAvision segmentation to reduce noise and improve quantification accuracy within each cluster. c, The segmentation of computed tomography images. By leveraging different annotated labels, structures with similar gray intensity values were successfully separated into distinct anatomical regions. d, The segmentation of MRI images enables visualization of the anatomy of the brain and spinal cord27,59,66. Part b adapted with permission from ref. 27, Elsevier. Part c adapted with permission from ref. 66, Elsevier.
First, we quantified microanatomical structures in liver and lung histology from a mouse model of pancreatic cancer (Fig. 14a). We demonstrate that CODAvision can be used to quantify metastatic burden and organ composition. This histological dataset was generated through an in vivo experiment originally described in the cited work59. We segmented six structures in the mouse lungs and seven structures in the mouse liver and achieved >90% overall accuracy and >85% per-class precision and recall in each model. With this segmentation, we demonstrated that in this mouse experiment, the metastases were larger, more numerous and more solid in the lung compared with the liver. While treatment of the mice with (diphtheria toxin) DT + anti-CD8 appeared to inhibit lung metastases, it did not affect the liver metastatic burden. The pipeline simultaneously analyzed the composition of other key tissues (bronchioles, alveoli, stroma and vasculature in the lungs, and hepatocytes, bile duct, stroma, fat and vasculature in the liver), demonstrating its utility for comprehensive compositional analysis in preclinical models. We present additional analysis of mouse lung histology in Extended Data Fig. 4 using example Dataset 1, where CODAvision was applied to quantify metastatic burden across 55 H&E-stained lung tissue sections from mice injected with MDA-MB-231 breast cancer cells60.
Second, we segmented cell types in human pancreas histology to deconvolute spot-based 10× Genomics Visium spatial transcriptomic data (Fig. 14b). Deconvolution of spatial transcriptomics data has emerged as a vital process for honing in on gene expression signatures of target cells that may make up a small fraction of a sample27,28,61–64. This dataset and the computational method for deconvolution using segmentation results, were originally described in the cited work27. We segmented nine structures of pancreatic microanatomy, including normal pancreatic ducts and pancreatic precancers, achieving >90% overall accuracy and >80% per-class precision and recall. The cell type labels at the coordinates of each 55-μm radius Visium spot were extracted and used to assess the cellular purity of each spot. This deconvolution allowed us to clearly determine the gene expression of pancreatic precancerous cells and to eliminate confounding gene expression signatures contributed by non-neoplastic cells within the same Visium spot.
Third, to demonstrate CODAvision’s broad applicability to biological images beyond histology, we segmented functional components of extracorporeal membrane oxygenation (ECMO) membranes in computed tomography (Fig. 14c). ECMO is a device used in critical care medicine to provide blood oxygenation for patients suffering from conditions that impact lung function such as coronavirus disease 2019 (ref. 65). The ECMO membrane is made up of five distinct fibrous layers contained within either a square or cylindrical exterior casing. Blood is perfused from the patient and through the multiple matrix design layers and the hollow fiber layers of the ECMO where gas exchange occurs before the blood is returned to the patient66. The dataset shown here was originally presented in the cited work67, where computed tomography images of used ECMO membranes were analyzed to determine which of the five fibrous layers contained the highest composition of blood clots (a common cause of ECMO failure). Owing to the highly specialized nature of the desired segmentation, more common methods for computed tomography segmentation were insufficient. CODAvision can be used to automate segmentation across every computed tomography slice, ensuring precise and consistent identification of layers and clot formations. We trained a model to detect seven structures in the ECMO device, including the five fibrous layers and blood clots, achieving >90% overall accuracy and >80% per-class precision and recall. After training, we applied the segmentation model to classify the 985 serial images making up the 3D computed tomography data and quantified the composition of blood clots in each of the fibrous layers, identifying layer 2 as a region with substantially higher percentage of clots. This analysis, possible only through segmentation of the subtle textures that define the distinct layers, has the potential to enable future design of more efficient ECMO devices that are less susceptible to blockage.
Finally, to further highlight the versatility of CODAvision, we applied our method to an MRI dataset of the human brain (Fig. 14d). Here, we desired to rapidly construct a highly specific model to distinguish the anatomical components of the brain, including white matter, grey matter, cerebellum and non-brain structures, including human eyes. We annotated these structures and trained a segmentation model with CODAvision, achieving >90% overall accuracy and >80% per-class precision and recall. After model training, we applied the segmentation model to the 157 serial images making up the 3D MRI data. Using the cited method68, we transformed the segmented .tif images outputted by CODAvision into a 3D printable .stl file. This enabled us to 3D print an anatomically correct model of the brain.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Extended Data
Extended Data Fig. 1 |. Classification of murine lung histology using CODAvision and QuPath.

A dataset of mouse lung histology was generated, and anatomical structures were annotated in these images using both Aperio ImageScope and QuPath: alveoli, bronchioles, stroma, cancer metastases, vasculature, and background. One image was used for independent model testing, with all others used for training. CODAvision performed equivalently (>90% overall accuracy and >85% per-class precision and recall) for annotations generated in either Aperio ImageScope or QuPath. QuPath pixel classification performed significantly worse using both annotation datasets, with ~33% overall accuracy and extensive over-calling of cancer metastases in the testing image.
Extended Data Fig. 2 |. Comparison of CODAvision to 3D Slicer for MRI segmentation.

a. CODAvision achieves robust performance in segmentation of MPRAGE sagittal brain MRI while 3D slicer fails to provide smooth boundaries between the tissues segmented. Arrows point at some of the misclassifications in the 3D Slicer seed-based segmentation compared with CODAvision b. 3D visualization of CODAvision segmentation. c. 3D visualization of 3D slicer segmentation.
Extended Data Fig. 3 |. Segmentation of MPRAGE sagittal brain MRI using MONAI in 3D Slicer.

(Top) Raw inference of 132 labels obtained with thewholeBrainSeg_Large_UNEST_segmentation MONAI bundle model. (Bottom) Result of the wholeBrainSeg_Large_UNEST_segmentation after manual refinement using the Segment Editor. Arrows point at some examples of the misclassification of this pretrained model.
Extended Data Fig. 4 |. Quantification of metastatic burden in mouse lung histology using CODAvision.

a. Representative semantic segmentation results comparing lung histology from a mouse used in a control arm (top) to histology from a mouse used in the experimental arm (bottom) of an in vivo experiment. b. Sample tissue composition analysis with metastases object count for sample a (bottom). c. Comparison of metastatic burden occupied in the lung across three experimental conditions, showing CODAvision quantification (left) versus a conventional ImageJ-based analysis (right). These results show a validations of CODAvision’s capability to quantify metastatic burden, we analyzed 55 H&E-stained lung tissue sections from mice injected with MDA-MB-231 breast cancer cells. Using a DeepLabv3+ architecture with a ResNet50 backbone (precision and recall >90% for all tissue types), we quantified the metastatic coverage, which revealed 69% and 52% metastatic burden (overall percentage of lung tissue) for scrambled control and wild-type cells, respectively (panel c, left). Part c adapted with permission from ref. 60, AAAS.
Extended Data Fig. 5 |. Segmentation of an independent pancreatic tissue slide using a CODAvision pretrained model.

(Top) Model trained on pancreatic tissue slides from Johns Hopkins Hospital. (Bottom) Inference on an independent pancreatic tissue slide obtained from the GTEx portal.
Supplementary Material
Supplementary information The online version contains supplementary material available at https://doi.org/10.1038/s41596-026-01404-3.
Key points.
The protocol details best practices for optimizing training datasets, parameterizing an intuitive user interface for model configuration and architecture selection, and automatically generating model performance reports and quantitative data analyses.
CODAvision is a segmentation framework that provides a customizable workflow for manual annotations to be parameterized and automatically converted into optimized training tiles, enabling users with limited computational experience to efficiently generate high-quality trainable datasets.
Acknowledgements
We thank the following sources of support: The Johns Hopkins University Data Science and Artificial Intelligence Program (grant no. U54CA268083), the National Institutes of Health (grant no. P51 OD011092), and the Break Through Cancer Data Science TeamLab. This work was partially supported by the Ministerio de Ciencia, Innovación y Universidades, Agencia Estatal de Investigación, MCIN/AEI/10.13039/501100011033, under grant no. PID2023-152631OB-I00, and co-financed by the European Regional Development Fund.
Footnotes
Competing interests
A.L.K., R.H.H., D.W. and L.D.W. are co-inventors on US patent application 18/572352 (Computational techniques for three-dimensional reconstruction and multi-labeling of serially sectioned tissue). The other authors declare no competing interests.
Extended data is available for this paper at https://doi.org/10.1038/s41596-026-01404-3.
Peer review information Nature Protocols thanks Rong Fan and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.
Data availability
Sample datasets and data underlying Figs. 2–4 and supplementary figures are publicly available through the Johns Hopkins Research Data Repository at https://doi.org/10.7281/T1W3JEFA (V2)69. The repository also includes a demo dataset of murine histology and the expected outputs following processing with CODAvision. Owing to patient confidentiality, the pancreas, skin and liver tissue samples used in the snapshot workflows are available upon request. The data used for the analyses in Extended Data Fig. 5 were obtained from the GTEx Portal on 30 October 2025 and from dbGaP (accession no. phs000424.v8.p2). Mouse histology used for Fig. 14a was analyzed from an in vivo experiment originally described in work conducted by de Groot et al. and is available upon request30. Pancreatic histology and sequencing data used for Fig. 14b are available upon request from the study by A.T.F.B. et al.27. Data corresponding to ECMO images segmented in Fig. 14c are available upon request from the study by J.S.H.W. et al.66. Human brain MRI images used in Fig. 14d are available upon request.
Code availability
The CODAvision source code is available via GitHub at https://github.com/Kiemen-Lab/CODAvision. A frozen version of the code used in this study is archived at the Johns Hopkins Research Data Repository at https://doi.org/10.7281/T1W3JEFA (ref. 69). The software is distributed under the MIT License, permitting free use, modification and distribution with appropriate attribution.
References
- 1.Chan HP, Samala RK, Hadjiiski LM & Zhou C. Deep learning in medical image analysis. Adv. Exp. Med. Biol. 1213, 3–21 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Lundervold AS & Lundervold A. An overview of deep learning in medical imaging focusing on MRI. Z. Med. Phys. 29, 102–127 (2019). [DOI] [PubMed] [Google Scholar]
- 3.Salto-Tellez M, Maxwell P & Hamilton P. Artificial intelligence—the third revolution in pathology. Histopathology 74, 372–376 (2019). [DOI] [PubMed] [Google Scholar]
- 4.Bera K, Schalper KA, Rimm DL, Velcheti V & Madabhushi A. Artificial intelligence in digital pathology—new tools for diagnosis and precision oncology. Nat. Rev. Clin. Oncol. 16, 703–715 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Paul J et al. Digital transformation: a multidisciplinary perspective and future research agenda. Int. J. Consum. Stud. 48, e13015 (2024). [Google Scholar]
- 6.Echle A et al. Deep learning in cancer pathology: a new generation of clinical biomarkers. Brit. J. Cancer 124, 686–696 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ghahremani P et al. Deep learning-inferred multiplex immunofluorescence for immunohistochemical image quantification. Nat. Mach. Intell. 4, 401–412 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zhang D et al. Inferring super-resolution tissue architecture by integrating spatial transcriptomics with histology. Nat. Biotechnol. 42, 1372–1377 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lu MY et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng. 5, 555–570 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xie W et al. Prostate cancer risk stratification via nondestructive 3D pathology with deep learning-assisted gland analysis. Cancer Res. 82, 334–345 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Phillip JM, Han KS, Chen WC, Wirtz D & Wu PH. A robust unsupervised machine-learning method to quantify the morphological heterogeneity of cells and nuclei. Nat. Protoc. 16, 754–774 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Chen LC, Papandreou G, Kokkinos I, Murphy K & Yuille AL. DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 40, 834–848 (2018). [DOI] [PubMed] [Google Scholar]
- 13.Liu Z et al. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. 2021 IEEE/CVF International Conference on Computer Vision (ICCV 2021) 9992–10002 10.1109/Iccv48922.2021.00986 (2021). [DOI] [Google Scholar]
- 14.Strudel R, Garcia R, Laptev I & Schmid C. Segmenter: transformer for semantic segmentation. In 2021 IEEE/CVF International Conference on Computer Vision 7242–7252 10.1109/Iccv48922.2021.00717 (2021). [DOI] [Google Scholar]
- 15.Belevich I & Jokitalo E. DeepMIB: User-friendly and open-source software for training of deep learning network for biological image segmentation. PLoS Comput. Biol. 17, e1008374 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Gomez-de-Mariscal E et al. DeepImageJ: a user-friendly environment to run deep learning models in ImageJ. Nat. Methods 18, 1192–1195 (2021). [DOI] [PubMed] [Google Scholar]
- 17.Lutnick B et al. A user-friendly tool for cloud-based whole slide image segmentation with examples from renal histopathology. Commun. Med. 2, 105 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ghahremani P, Marino J, Dodds R & Nadeem S. DeepLIIF: an online platform for quantification of clinical pathology slides. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition 21367–21373 10.1109/Cvpr52688.2022.02071 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Muller A et al. Modular segmentation, spatial analysis and visualization of volume electron microscopy datasets. Nat. Protoc. 19, 1436–1466 (2024). [DOI] [PubMed] [Google Scholar]
- 20.Kiemen AL et al. CODA: quantitative 3D reconstruction of large tissues at cellular resolution. Nat. Methods 19, 1490–1499 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kiemen AL et al. PanIN or IPMN? Redefining lesion size in three dimensions. Am. J. Surg. Pathol. 48, 839–845 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Johnston AC et al. Engineering self-propelled tumor-infiltrating CAR T cells using synthetic velocity receptors. Preprint at biorxiv 10.1101/2023.12.13.571595 (2024). [DOI] [Google Scholar]
- 23.Dequiedt L et al. Three-dimensional reconstruction of fetal rhesus macaque kidneys at single-cell resolution reveals complex inter-relation of structures. Preprint at bioRxiv 10.1101/2023.12.07.570622 (2023). [DOI] [Google Scholar]
- 24.Kiemen AL et al. Tissue clearing and 3D reconstruction of digitized, serially sectioned slides provide novel insights into pancreatic cancer. Med 4, 75–91 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Steele NG et al. Inhibition of Hedgehog signaling alters fibroblast composition in pancreatic cancer. Clin. Cancer Res. 27, 2023–2037 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Joshi S et al. InterpolAI: deep learning-based optical flow interpolation and restoration of biomedical images for improved 3D tissue mapping. Nat. Methods 22, 1556–1567 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Bell ATF et al. PanIN and CAF transitions in pancreatic carcinogenesis revealed with spatial data integration. Cell Syst 15, 753–769.e5 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Deshpande A et al. Uncovering the spatial landscape of molecular interactions within the tumor microenvironment through latent spaces. Cell Syst. 14, 285–301 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Sneider A et al. Deep learning identification of stiffness markers in breast cancer. Biomaterials 285, 121540 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.de Groot AE et al. Characterization of tumor-associated macrophages in prostate cancer transgenic mouse models. Prostate 81, 629–647 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Kiemen AL et al. Intraparenchymal metastases as a cause for local recurrence of pancreatic cancer. Histopathology 82, 504–506 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Crawford AJ et al. Combined assembloid modeling and 3D whole-organ mapping captures the microanatomy and function of the human fallopian tube. Sci. Adv. 10, eadp6285 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Bai B et al. Label-free virtual HER2 immunohistochemical staining of breast tissue using deep learning. BME Front. 2022, 9786242 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Xue Y et al. Mechanical tension mobilizes Lgr6+ epidermal stem cells to drive skin growth. Sci. Adv. 8, eabl8698 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Yang H et al. Engineered bispecific antibodies targeting the interleukin-6 and −8 receptors potently inhibit cancer cell migration and tumor metastasis. Mol. Ther. 30, 3430–3449 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Forjaz A et al. Three-dimensional assessments are necessary to determine the true, spatially resolved composition of tissues. Cell Rep. Methods 5, 101075 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Forjaz A et al. 3D multi-omic mapping of whole nondiseased human fallopian tubes at cellular resolution reveals a large incidence of ovarian cancer precursors. Preprint at bioRxiv 10.1101/2025.09.21.677628 (2025). [DOI] [Google Scholar]
- 38.Gensbigler PA, Foster W, Kiemen AL & Bever GS. An optimized, high-throughput workflow for the collection, processing, and visualization of histology data in comparative vertebrate morphogenesis. Dev. Dyn 10.1002/dvdy.70092 (2025). [DOI] [PubMed] [Google Scholar]
- 39.Than MT et al. Cancer interception with KRAS inhibitors in preclinical models of pancreatic ductal adenocarcinoma. Science 391, 1161–1166 (2026). [DOI] [PubMed] [Google Scholar]
- 40.O’Brien J et al. Skin keratinocyte-derived SIRT1 and BDNF modulate mechanical allodynia in mouse models of diabetic neuropathy. Brain 147, 3471–3486 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Braxton AM et al. 3D genomic mapping reveals multifocality of human pancreatic precancers. Nature 629, 679–687 (2024). [DOI] [PubMed] [Google Scholar]
- 42.Sidiropoulos DN et al. Machine learning integrating spatial omics uncovers humoral immunity patterns in intratumoral tertiary lymphoid structures in pancreatic cancer pathologic responders. Cancer Res. 10.1158/1538-7445.Am2024-1159 (2024). [DOI] [Google Scholar]
- 43.Kiemen AL et al. Three-dimensional immune atlas of pancreatic cancer precursor lesions reveals large inter- and intra-lesion heterogeneity. Cancer Res. 10.1158/1538-7445.Am2024-1206 (2024). [DOI] [Google Scholar]
- 44.Lee MH et al. Multi-compartment tumor organoids. Mater. Today 61, 104–116 (2022). [Google Scholar]
- 45.English IA et al. Myc and Kras cooperate in adult acinar cells to drive phenotypic heterogeneity, metastasis, and therapeutic resistance in a novel pancreatic cancer mouse model. Preprint at bioRxiv 10.1101/2025.07.14.664767 (2025). [DOI] [Google Scholar]
- 46.Montezuma D et al. Annotation practices in computational pathology: a European Society of Digital and Integrative Pathology (ESDIP) survey study. Lab. Invest. 105, 102203 (2024). [DOI] [PubMed] [Google Scholar]
- 47.Liao YH, Kar A & Fidler S. Towards good practices for efficiently annotating large-scale image classification datasets. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition 4348–4357 10.1109/Cvpr46437.2021.00433 (2021). [DOI] [Google Scholar]
- 48.Kirillov A et al. Segment anything. In IEEE Int. Conf. Comp. Vis. 3992–4003 10.1109/Iccv51070.2023.00371 (2023). [DOI] [Google Scholar]
- 49.Moor M et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023). [DOI] [PubMed] [Google Scholar]
- 50.Zhang K et al. A generalist vision-language foundation model for diverse biomedical tasks. Nat. Med. 30, 3129–3141 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Chen RJ et al. Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition 16123–16134 10.1109/Cvpr52688.2022.01567 (2022). [DOI] [Google Scholar]
- 52.Bankhead P et al. QuPath: open source software for digital pathology image analysis. Sci. Rep. 7, 16878 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Isensee F, Jaeger PF, Kohl SAA, Petersen J & Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods 18, 203–211 (2021). [DOI] [PubMed] [Google Scholar]
- 54.Fedorov A et al. 3D Slicer as an image computing platform for the Quantitative Imaging Network. Magn. Reson. Imaging 30, 1323–1341 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Stringer C, Wang T, Michaelos M & Pachitariu M. Cellpose: a generalist algorithm for cellular segmentation. Nat. Methods 18, 100–106 (2021). [DOI] [PubMed] [Google Scholar]
- 56.Bannon D et al. DeepCell Kiosk: scaling deep learning-enabled cellular image analysis with Kubernetes. Nat. Methods 18, 43–45 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Goode A, Gilbert B, Harkes J, Jukic D & Satyanarayanan M. OpenSlide: a vendor-neutral software foundation for digital pathology. J. Path. Inform. 4, 27 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Ronneberger O, Fischer P & Brox T. U-net: convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention 234–241 (Springer, 2015). [Google Scholar]
- 59.Zhang Y et al. Myeloid cells are required for PD-1/PD-L1 checkpoint activation and the establishment of an immunosuppressive environment in pancreatic cancer. Gut 66, 124–136 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Nair PR et al. MLL1 regulates cytokine-driven cell migration and metastasis. Sci. Adv. 10, eadk0785 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Li H et al. A comprehensive benchmarking with practical guidelines for cellular deconvolution of spatial transcriptomics. Nat. Commun. 14, 1548 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Dong R & Yuan GC. SpatialDWLS: accurate deconvolution of spatial transcriptomic data. Genome Biol. 22, 145 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Zhou Z, Zhong Y, Zhang Z & Ren X. Spatial transcriptomics deconvolution at single-cell resolution using Redeconve. Nat. Commun. 14, 7930 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Forjaz A et al. PIVOT: an open-source tool for multi-omic spatial data registration. Preprint at bioRxiv 10.1101/2025.06.08.658506 (2025). [DOI] [Google Scholar]
- 65.Olson SR et al. Thrombosis and bleeding in extracorporeal membrane oxygenation (ECMO) without anticoagulation: a systematic review. ASAIO J. 67, 290–296 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Wang JSH et al. Multimodality quantification of thrombus deposition in extracorporeal membrane oxygenation (ECMO): correlating oxygenator computed tomography imaging, electron microscopy and histology to clinical outcomes. Blood 142, 1285 (2023). [Google Scholar]
- 67.Wang JSH et al. Development of a method for visualizing and quantifying thrombus formation in extracorporeal membrane oxygenators. Cell. Mol. Bioeng. 18, 197–209 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Kiemen AL et al. High-resolution 3D printing of pancreatic ductal microanatomy enabled by serial histology. Adv. Mater. Technol. 9, ARTN 2301837 (2024). [Google Scholar]
- 69.Matos-Romero V & Kiemen AL. Data associated with the publication: CODAvision: best practices and a user-friendly interface for rapid, customizable segmentation of medical images. Johns Hopkins Research Data Repository 10.7281/T1W3JEFA (2025). [DOI] [Google Scholar]
Key references
- Kiemen AL et al. Nat. Methods 19, 1490–1499 (2022): 10.1038/s41592-022-01650-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Braxton AM et al. Nature 629, 679–687 (2024): 10.1038/s41586-024-07359-3 [DOI] [PubMed] [Google Scholar]
- Bell ATF et al. Cell Syst. 15, 753–769.e5 (2024): 10.1016/j.cels.2024.07.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang JSH et al. Cell. Mol. Bioeng. 18, 197–209 (2025): 10.1007/s12195-025-00847-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Than MT et al. Science 391, 1161–1166 (2026): 10.1126/science.aec7929 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Sample datasets and data underlying Figs. 2–4 and supplementary figures are publicly available through the Johns Hopkins Research Data Repository at https://doi.org/10.7281/T1W3JEFA (V2)69. The repository also includes a demo dataset of murine histology and the expected outputs following processing with CODAvision. Owing to patient confidentiality, the pancreas, skin and liver tissue samples used in the snapshot workflows are available upon request. The data used for the analyses in Extended Data Fig. 5 were obtained from the GTEx Portal on 30 October 2025 and from dbGaP (accession no. phs000424.v8.p2). Mouse histology used for Fig. 14a was analyzed from an in vivo experiment originally described in work conducted by de Groot et al. and is available upon request30. Pancreatic histology and sequencing data used for Fig. 14b are available upon request from the study by A.T.F.B. et al.27. Data corresponding to ECMO images segmented in Fig. 14c are available upon request from the study by J.S.H.W. et al.66. Human brain MRI images used in Fig. 14d are available upon request.
The CODAvision source code is available via GitHub at https://github.com/Kiemen-Lab/CODAvision. A frozen version of the code used in this study is archived at the Johns Hopkins Research Data Repository at https://doi.org/10.7281/T1W3JEFA (ref. 69). The software is distributed under the MIT License, permitting free use, modification and distribution with appropriate attribution.
