Abstract
Multi-modal learning, in which diverse data types are integrated and analyzed together, has become a central area of research in artificial intelligence, driving major advances in a wide range of domains. However, in many practical situations, certain modalities or variables may be missing for part of the samples, leading to a limited performance or failure of conventional methods. This has given a rise to the field of multi-modal learning with incomplete data, an area that has grown rapidly due to its broad real-world applications. Despite this, the community still lacks standardized tools to effectively handle incomplete multi-modal data. To fill this gap, we developed iMML, a unified, user-friendly Python package with versatile methods designed for integrating, processing, and analyzing incomplete multi-modal data. Successful use cases in biomedicine, text analysis, and computer vision for diverse machine learning tasks show the potency of iMML for making the best use of modern datasets in complex real-world applications. The iMML package is available at https://github.com/ocbe-uio/imml with an extensive documentation at https://imml.readthedocs.io/.
Subject terms: Machine learning, Scientific data
Incomplete multi-modal datasets pose a major challenge for real-world machine learning applications. Here, authors present iMML, a unified open-source Python package designed to analyze and integrate incomplete multi-modal datasets for diverse machine learning tasks.
Introduction
Multi-modal learning (MML) has emerged as a critical field in artificial intelligence. MML reflects the human ability to combine multiple inputs, such as sight, sound, movement, touch, and smell, hence leading to a richer understanding of the environment. Multi-modal approaches can identify underlying patterns and interactions between modalities that may be missed by methods relying on a single modality1. As a result, MML is having a profound impact across numerous domains, such as autonomous systems, healthcare, multimedia analysis, and natural language processing2–6.
However, multi-modal datasets are often incomplete due to various reasons, including measurement failures, hardware restrictions, privacy limitations, high costs of data collection, transmission problems, or the lack of standardized formats7. The problem of incomplete multi-modal data (IMMD) can arise at any stage from data collection to deployment, and it can significantly impact model performance. For instance, The Cancer Genome Atlas (TCGA), a widely used resource for cancer research, contains numerous missing omics modalities for a subset of the patients8. In traffic analysis, missing sensor readings caused by natural and human factors can significantly affect decision-making9. In human-computer interaction, missing images or audio can result from tracking failures due to occlusion10. These real-world examples highlight the complexity of working with multi-modal data across diverse environments.
Learning from IMMD has seen important growth last years (Fig. 1a)11. Numerous recent papers have identified the field as a promising area for future research12–17. Adapting existing approaches to handle IMMD has also been widely suggested as future work18–20. Despite this progress, several limitations still persist. First, most published methods do not provide open-source implementations (Fig. 1b, c). Even when code is released, it is typically restricted to reproducing the results reported in the original paper, rather than being designed as a reusable tool for the broader community. Moreover, the landscape of the few available methods is fragmented, largely due to the diversity of use cases and data modalities, which complicates both their application and benchmarking. Systematic use and comparison of the current methods are further hindered by practical challenges, such as incompatible input data formats and conflicting software dependencies. As a result of the lack of robust and standardized tools, researchers frequently face challenges in choosing a practical method and invest considerable efforts into reconciling codebases, rather than addressing the core scientific questions.
Fig. 1. Recent growth in the field of multi-modal learning with incomplete data.

a The number of publications is increasing almost exponentially every year. b Only a third of publications shared open-source code. In terms of the programming language, MATLAB dominated the early years of the field, but Python has since surpassed MATLAB (the proportion is indicated in the middle of the bars). The apparent decline in 2025 is likely due to authors releasing their code after the publication date. c The tasks addressed by publications with open-source code, with an indication of whether the iMML package supports those tasks. Other tasks include regression (4), data alignment (3), generation (3), survival (2), tracking (2), explainability (1), and question-answering (1). For more information on this search, see Supplementary Methods.
When researchers lack appropriate tools to handle IMMD, they tend to rely on three basic approaches21: (i) selecting only the samples with all modalities available, (ii) using a single modality with most data available, or (iii) applying a simple imputation method. The first approach leads to significant data loss; for example, various combinations of omics data were tested for developing a survival prediction model, but ultimately the model used only patients with all modalities, which excluded 47% of the patients22. The second approach sacrifices the rich, complementary information that multi-modal data provides, potentially reducing its performance benefits1. Lastly, naive imputation methods can introduce noise, and there are no universally well-performing imputation approaches for all the modalities and domains23. Consequently, researchers are seeking more sophisticated methods to deal with missing modalities and hence make the most of their IMMD.
To address this gap, we developed iMML, a Python package designed for MML with incomplete data (Fig. 2). The key features of this package are:
Coverage: iMML offers more than 25 methods for integrating, processing, and analyzing incomplete multi-modal datasets implemented as a single, user-friendly interface to facilitate adoption by a wide community of users. The package includes extensive technical testing to ensure robustness.
Comprehensive: designed to be compatible with widely-used machine learning and data analysis tools, allowing use with minimal programming effort. An extensive documentation enables end-users to apply its functionality effectively.
Extensible: iMML provides a unified framework where researchers can contribute and integrate new approaches, serving as a community platform for hosting new algorithms and methods.
Fig. 2. Overview of iMML.

iMML is a Python package that provides a robust toolset for integrating, processing, and analyzing incomplete multi-modal datasets to support a wide range of machine learning tasks. Starting with a dataset containing N samples (SN) with K modalities (MK), iMML effectively handles missing data for classification, clustering, data retrieval, imputation and amputation, feature selection, feature extraction, and data exploration, hence enabling efficient analysis of partially observed samples.
Results
iMML integrates methods originally developed across diverse application domains, resulting in substantial heterogeneity in terms of implementation languages and software design. To accommodate this diversity, iMML was engineered as a versatile and robust platform that supports a broad research community and a wide range of use cases (Fig. 3).
Fig. 3. Example pipeline in a multi-modal clinical application.

The input consists of a multi-modal dataset comprising clinical variables in tabular form, a histopathological whole-slide image (WSI), and a text report written by a clinician describing the image (such as the PRAD dataset82). The user first loads the multi-modal inputs using standard frameworks such as pandas, NumPy, or PyTorch. As not all samples contain all modalities, the availability of patients is examined as a function of the modalities present. To further explore the data, the user assesses whether the information captured in the images is redundant with the clinical variables, which could potentially simplify the learning problem. Any of the algorithms implemented in iMML is then trained and saved using an interface consistent with widely adopted machine learning frameworks, such as scikit-learn and Lightning AI, ensuring ease of use for both experienced and new users. Finally, the trained model is applied to unseen samples to generate predictions and evaluate prediction performance.
To demonstrate the utility of the iMML package, we present six illustrative use cases that cover different domains, including biomedicine, text analysis, and computer vision. Although not exhaustive, these examples demonstrate representative real-world scenarios and provide insight into the broader capabilities and potential of iMML to handle complex, incomplete multi-modal datasets across diverse domains.
Comparison with existing libraries
Comparing iMML with the other popular open-source packages, summarized in Table 1, we note that none of the existing libraries addresses the challenge of MML with incomplete data:
Imputation and amputation. HyperImpute, Scikit-learn, and mdatagen offer a variety of imputation methods. The popular R mice package is widely used for both imputation and amputation. Another popular library for data amputation is pyampute. While these libraries are useful and the common choices for imputing missing data, they are limited to uni-modal data and do not support multi-modal contexts.
Multi-modal learning. mvlearn offers a range of algorithms for different tasks, following a scikit-learn-style interface. However, we found it is no longer maintained (last commit in 2022). Also, scikit-multimodallearn follows the same API design, although it is limited to classification tasks. In both cases, they only support fully observed multi-tabular data. MultiZoo/MultiBench is a repository (and therefore not an actual package published in the Package Index, PyPI), which includes a large diversity of multi-modal datasets and fusion paradigms, optimization objectives, and training approaches. TorchMultimodal is another interesting library, although it includes only deep learning algorithms. The main problems we found were limited documentation, the most recent supported Python version is 3.9, and it is not maintained. None of the packages addresses learning with incomplete data.
Multi-modal learning with incomplete data. The only notable repository in this space was developed in MATLAB, where the authors tried to provide a framework to unify incomplete multi-view clustering. The implementation in MATLAB limits the accessibility and ease of use of the framework, which is why we translated the clustering algorithms available in this repository. Additionally, its scope is restricted to unsupervised clustering.
Table 1.
Use cases supported by iMML and other open source libraries (if not otherwise specified in the table, the package is implemented in Python)
| Package | Imputation | Amputation | Machine learning | ||||
|---|---|---|---|---|---|---|---|
| UM | MM | UM | MM | UM | MM | IMM | |
| iMML | ✓ | ✓ | ✓ | ✓ | |||
| Survey_IMC*73 (MATLAB) | ✓(Cluster) | ✓(Cluster) | |||||
| MultiZoo*74 (2023) | ✓ | ||||||
| TorchMultimodal (2022) | ✓ | ||||||
| scikit-multimodallearn75 (2022) | ✓ | ||||||
| HyperImpute76 (2022) | ✓ | ||||||
| pyampute77,78 (2022) | ✓ | ||||||
| mvlearn79 (2021) | ✓ | ||||||
| mdatagen80 (2019) | ✓ | ||||||
| mice81 (R) (2011) | ✓ | ✓ | |||||
| Scikit-learn60 (2011) | ✓ | ✓ | |||||
Libraries are sorted by decreasing year of publication. Comparisons are based on the software versions available at the time of writing. * denotes repositories rather than packages.
UM uni-modal, MM multi-modal, IMM incomplete multi-modal.
In summary, none of the existing libraries addresses the challenges posed by incomplete multi-modal datasets. This is where the iMML package fills a critical gap by providing researchers and the machine learning community with a toolset and user-friendly environment to work effectively with incomplete multi-modal datasets.
Implementation analysis
Methods in iMML were originally developed in several programming languages (Python, MATLAB, and R). We translated most algorithms into native Python to offer pure Python code and reduce dependencies. These translations not only improved transparency but also enhanced performance (Table 2, Supplementary Table 1 and Source Data 1); notably, most algorithms experienced substantial speed improvements (1–10×) and reduced memory usage in Python. Both the translations and the original code provided almost identical results, with p = 1 for all methods (Nemenyi-Wilcoxon-Wilcox test), indicating the new Python implementations do not deviate statistically from the original MATLAB/R codes (Supplementary Figs. 1 and 2).
Table 2.
Computing performance comparison between the original and the new Python implementations
| Algorithm | Speed-up factor (×) | Memory reduction factor (×) |
|---|---|---|
| LFIMVC | 9.33 (3.08, 16.55) | 1.82 (1.09, 3.20) |
| EEIMVC | 8.45 (1.78, 16.23) | 1.54 (1.07, 2.35) |
| NEMO | 5.61 (2.45, 10.81) | 2.02 (1.01, 4.71) |
| SIMCADC | 5.47 (2.19, 10.64) | 1.63 (1.10, 2.26) |
| IMSR | 3.69 (1.11, 13.86) | 1.21 (0.44, 2.23) |
| DAIMC | 0.97 (0.24, 2.09) | 1.50 (1.03, 2.19) |
For each algorithm, the average speed-up and memory gain over five datasets and 50 repetitions are reported, with minimum and maximum values indicated in parentheses. Values larger than 1 indicate improved performance of the Python implementation (e.g., a value of 2 corresponds to a two-fold speed-up or a two-fold reduction in memory usage). The Python implementations show clear speed improvements and lower memory consumption in most cases.
We also offer the original implementations through the wrappers Oct2Py for MATLAB and rpy2 for R. This allows users to run the algorithms in Python while accessing other programming languages (via wrappers) or entirely in Python.
To assess scalability within the iMML framework, we evaluated two state-of-the-art classification methods, MUSE and M3Care, under varying dataset sizes, training configurations, and model parameters. The results show that the number of samples had the strongest impact on peak memory consumption for both methods, followed by the number of modalities (Supplementary Fig. 3 and Source Data 2). In contrast, runtime was most strongly influenced by the number of training epochs.
Exploring an incomplete multi-omics dataset
Multi-omics data constitute a cornerstone of precision medicine, and rigorous quality assessment is a critical first step in their analysis. To support this process, iMML provides a suite of visualization and exploratory analysis tools tailored to incomplete multi-modal datasets. As an illustrative example, we analyzed pancreatic cancer patients from TCGA, integrating copy number alterations (CNA), RNA sequencing (gene expression; GEx), proteomics, and mutation data24.
A summary of data incompleteness reveals that CNA is fully observed across all patients, whereas mutation data exhibit missing variables for multiple samples (Fig. 4a). Moreover, the number of patients available varies substantially depending on the combination of modalities considered, ranging from 177 patients when using CNA and gene expression to only 100 patients when all four omics layers are required (Fig. 4b). While related to classical UpSet plots, our visualization directly reports the number of samples available for each modality combination, rather than the union of set intersections. We found this representation to be more intuitive and better aligned with the practical needs of researchers working with incomplete multi-modal data.
Fig. 4. Use cases.

a, b Overview of available and missing data for pancreatic cancer patients with four omics layers from TCGA. c Quantification of their redundancy (R), uniqueness (U1 and U2), and synergy (S). d, e Amputation data patterns in a dataset with 10 samples and 4 modalities with 80% incomplete samples. The white areas represent missing data, while the shaded areas denote observed data across the samples. The column ("Factor") represents the (observed or unobserved) factor that explains the missing not at random pattern. f Classification performance of RAGPT and M3Care on the Food101 dataset using a stratified five-fold cross-validation across various missing rates. A linear model was fitted for each model and missingness pattern to estimate how the performance varies when using incomplete data. Results are averaged across folds, with 95% confidence intervals indicated. g Retrieved instances for an incomplete vision-language dataset (Food101) when using image-only (top) and text-only (bottom) as examples of target instances. h The clustering solutions provided by five clustering algorithms, as well as their respective baselines (where missing values were replaced with the feature-wise average), are compared to ground truth topics across various missing rates for a text dataset (BBCSport). The average of 50 repetitions with 95% confidence intervals is shown. i Performance of a classifier using extracted features, selected features, all features, and randomly selected features across varying missing rates on an incomplete nutrigenomic dataset. Results are averaged over 50 repetitions, with 95% confidence intervals shown. j Top selected features, k most influential features of the extracted features, and l modality importance with a 20% missing rate.
Beyond visualization, iMML enables quantification of the information contributed by each modality. Specifically, the package includes a module for estimating multi-modal statistics, such as redundancy, uniqueness, and synergy across modalities. Applied to the TCGA pancreatic cancer dataset, this analysis suggests that the combination of gene expression and proteomics is particularly informative for patient survival prediction, offering high values of synergy and low redundancy (Fig. 4c, Supplementary Table 2 and Source Data 3). To the best of our knowledge, no systematic benchmarking of the best combination of data modalities for survival prediction using these four omics data types exists to validate our approach. However, multiple proteogenomic studies in pancreatic cancer consistently show that proteomic data provides non-redundant and clinically relevant information beyond transcriptomics (median gene-wise correlation of 0.3225), and that integrating RNA and protein layers improves survival association26,27.
Together, these visualization, exploration and statistical modules provide essential tools for early-stage data exploration and decision-making, facilitating informed selection of modalities prior to applying more advanced modeling techniques.
Block-wise missing data generation for algorithm testing
Evaluation and benchmarking of new algorithms under diverse conditions is essential to ensure their robustness, added value, and generalizability. iMML simplifies this process by simulating block-wise IMMD. This so-called data amputation process allows for controlled testing of methods by generating missing data from various missingness patterns, thereby reflecting real-world scenarios, where different data modalities may be either partially observed or entirely missing. The availability of original data enables one to compare the results obtained on an incomplete multi-modal dataset when using various learning methods and algorithms.
As a simple example of a fully observed multi-modal dataset with 10 samples and 4 modalities, Fig. 4d, e illustrates the four block-wise missing patterns (MEM, PM, MCAR, MNAR) supported by iMML. Each pattern represents a practically relevant setting. For instance, the MNAR pattern simulates multi-center studies in which each site collects only a particular subset of modalities, whereas the PM pattern reflects situations in which some modalities are inexpensive or easy to acquire, while others require substantially greater effort or resources. The MEM pattern is particularly useful for evaluating model performance under extreme missingness conditions. Thus, while all the cases have the same number of complete and incomplete samples, each pattern represents a unique distribution of missing data across different modalities, helping researchers to assess the robustness of machine learning models in the presence of incomplete data.
Classification and retrieval of an incomplete vision-language dataset
Vision-language datasets are widely used in multiple fields, such as multimedia and healthcare, making their analysis one of the most common tasks in MML. As an illustrative example, we used the well-known Food101 dataset, an image-text dataset for food recognition28.
We trained the two vision-language classification algorithms currently available in iMML, namely RAGPT and M3Care. RAGPT was highly robust to missing modalities and consistently achieved the best performance across all evaluated conditions. M3Care also demonstrated strong robustness to incompleteness, although its overall predictive performance was lower (Fig. 4f, Supplementary Fig. 4 and Source Data 4). The multi-channel retriever required to train RAGPT also effectively identified relevant cross-modal instances (Fig. 4g).
This example illustrates how iMML enables state-of-the-art performance in classification and retrieval tasks29, even in the presence of significant modality incompleteness in vision-language datasets.
Clustering incomplete multi-modal text data
iMML supports clustering of IMMD, allowing users to perform unsupervised clustering without requiring complete data across all modalities. To showcase this capability, we used the BBCSport dataset, which consists of 116 sports news articles across five categories (athletics, cricket, football, rugby, and tennis), with four different views corresponding to different parts of the text30.
For this use case, we applied five algorithms from the iMML library, namely IMSCAGL, IMSR, LFIMVC, PIMVC, and SIMCADC, to cluster the dataset under varying percentages of incomplete samples and different missingness patterns. Missing modalities were simulated using the iMML ampute module. For comparison, we included baseline versions of each method, where the missing modalities were first imputed with mean values before clustering.
The results show that methods designed for incomplete data consistently outperform prior imputation across the tested settings (Fig. 4h, Supplementary Fig. 5 and Source Data 5). As expected, the clustering performance decreased when more incomplete samples were included. However, even with a high percentage of incomplete samples (>60%), several algorithms produced meaningful clustering results, highlighting their robustness in handling IMMD. Notably, IMSR performed best across most tested conditions, whereas PIMVC is a very competitive option at high missingness rates.
This example shows how iMML effectively supports clustering with IMMD, demonstrating the applicability of iMML for unsupervised clustering tasks in real-world multi-modal scenarios.
Imputation of block- and feature-wise incomplete multi-modal data in computer vision
When the learning algorithms cannot directly handle missing data, imputation methods become essential to allow their application. To demonstrate this task, we used the Statlog dataset, an image dataset containing 2310 instances across 7 outdoor image classes, where features (describing intensity, color, etc.) are organized into two views31.
We simulated block- and feature-wise missing data and then applied MOFA and DFMF for missing value imputation, and evaluated the imputation accuracy against real values. A baseline of mean imputation was also computed for comparison. After imputation, the modalities were concatenated and used as input for logistic regression to classify the instances. Both methods clearly outperformed the mean imputation baseline, both in terms of imputation and classification performance (Supplementary Figs. 6 and 7 and Source Data 6). In particular, MOFA achieved the best results across all conditions (average improvement in accuracy of 4–7% compared to the baseline for missing rates between 30 and 70%).
Since most MML algorithms do not accept missing data as input, this imputation step is essential for enabling their application in real-world multi-modal settings.
Feature extraction and selection on incomplete multi-modal biomedical data
High-dimensional datasets can severely impact machine learning projects, often degrading performance due to the curse of dimensionality, as well as the presence of correlated, noisy, or irrelevant features. These issues also increase computational demands and reduce model interpretability. Consequently, reducing the number of features is often critical. High-dimensional datasets are particularly common in fields like biomedical research, where the data collection is often not only complex but also costly. For this use case, we used the nutrimouse dataset, which comes from a mouse nutrigenomic study32. The dataset includes gene expression levels (120 variables) and concentrations of fatty acids (21 variables) measured in 40 mice, each labeled by two variables: genetic type (two classes) and diet (five classes).
As in previous examples, we simulated block- and feature-wise missing data to reflect real-world biomedical datasets. We applied the jNMF algorithm for feature extraction and feature selection, followed by a genetic classification task. For comparison, we also included baselines using randomly selected features and all available features (Supplementary Fig. 8 and Source Data 7). The extracted features demonstrated robustness, especially in scenarios with higher levels of missing data, where they surpassed the selected features (Fig. 4i). This was largely because the selected features required imputation post-selection, introducing noise and potential inaccuracies.
Beyond dimensionality reduction, feature selection is particularly valuable in biomarker discovery, where identifying the most relevant features is essential for understanding the underlying problem or task. As showed, the top features were originated from both gene expression and fatty acid measurements, with two genes and two fatty acid identified as the most important features, hence forming a multi-modal biomarker panel (Fig. 4j). The extracted features, representing aggregated versions of the original features, were further decomposed to gain insights into the model’s behavior (Fig. 4k). Quantifying the relative importance of the modalities revealed that gene expression contributed much more than fatty acid measurements. (Fig. 4l)
This case demonstrates the power of iMML for feature selection and extraction, key steps in biomedical research, where identifying biomarkers is crucial for personalized medicine. The reduced feature matrices can also be used for other downstream tasks, such as classification or regression, as also shown here, supporting further potential of the package for MML.
Discussion
Our vision is to establish iMML as the leading library for MML. In this first version of iMML, we focused on algorithms that operate on numerical representations, a decision that is both practical and strategic. Today, virtually all data modalities can be transformed into numeric vectorized forms, hence enabling a unified framework for their analysis. More importantly, this choice aligns with a growing need in the research community to support tabular data, which remains the dominant modality in scientific datasets and will also play a critical role in the future of MML33. To further broaden accessibility and encourage wide adoption, we have also integrated modality-specific methods, including language and vision-transformer-based architectures, ensuring that iMML accommodates both general-purpose and specialized methods and use cases.
iMML offers several important advantages in terms of technical robustness and reproducibility relative to the original implementations of the methods. Most original implementations were developed for a specific dataset and often restricted to a fixed number of modalities or data formats. We generalized these implementations to support arbitrary numbers of modalities and a broader range of data structures, substantially increasing their practical applicability. The framework allows end-users to control for stochasticity, so that repeated runs with the same experiment and conditions return exactly the same result, while many of the original methods lacked such a mechanism. This is crucial for reproducibility, because otherwise results cannot be reliably replicated. Furthermore, the code of the original implementations was intended to reproduce the result of the corresponding articles, not offering a tool that other users could easily use (except for MOFA or NEMO). This means that, as software ecosystems evolve, such code often becomes unusable due to dependency conflicts. Instead, the iMML framework uses version control, tags, and released versions through GitHub and the Python Package Index (PyPI), which means that each use of the methods is tied to a specific set of requirements, and only widely used and well-maintained libraries are required. Finally, in the case of deep learning models, iMML works with Lightning AI, a powerful deep learning framework that provides full control over the machine learning workflow and makes it easier to handle advanced features such as GPU parallelization, model quantization, or automatic learning-rate search.
The experimental case studies in the present work effectively demonstrate the usage of the package and illustrate its flexibility across a wide range of real-life applications. Given the broad range of applications involving multi-modal data, a large-scale and systematic benchmark of the various algorithms represents an important future research direction. Such a systematic study would enable the comprehensive comparison of algorithms across multiple datasets, missingness mechanisms, tasks, domains, and computational settings, providing an extremely useful resource for the community. We welcome practitioners, researchers, and the open-source community to contribute to the iMML project, and in doing so, helping us extend, refine, and benchmark the library for the community. Such a community-wide effort will make iMML more versatile, sustainable, powerful, and accessible to the machine learning community across many domains.
MML is becoming increasingly important, yet most existing algorithms assume fully observed data, an assumption that is often unrealistic in real-world scenarios. To fill this gap, we introduced iMML, an open-source library designed to address the challenges posed by incomplete data for MML. As a comprehensive toolkit with a consistent and user-friendly interface, iMML provides a unique resource for handling incomplete multi-modal data and offers a unified and extendable platform where researchers can contribute and integrate new approaches. Case examples covering a range of domains, including biomedicine, computer vision, and text analysis, demonstrate its versatility and robustness to support MML tasks in the context of incomplete data. Looking ahead, iMML has the potential to become a foundational tool for future advancements in MML, addressing emerging challenges and expanding the possibilities for machine learning in real-world applications.
Methods
The iMML library
Modular architecture
iMML consists of several modules, each addressing a particular task, including amputation, classification, clustering, data retrieval, exploration, imputation, feature selection, feature extraction, statistics, and visualization. It is designed so that the modules integrate in a seamless way with one another to build pipelines, hence providing a consistent and intuitive interface for end-users. The following sections provide an overview of the modules included in iMML, showcasing the versatility and depth of the library.
- Ampute. In this context, amputation refers to generating modality-wise missing data within otherwise complete multi-modal datasets. This module enables testing MML methods by simulating missing modalities based on four different missingness patterns (see Supplementary Methods for details):
- Missing completely at random (MCAR): missing modalities occur randomly.
- Missing not at random (MNAR): specific samples are missing in certain modalities due to known or unknown factors influencing the data missingness pattern.
- Mutually exclusive missing (MEM): incomplete samples have only one modality. It can be considered as an extreme case of MNAR. By default, all modalities contribute equally to the dataset, unlike other patterns where the informativeness of the modality affects its impact.
- Partial missing (PM): some of the modalities are fully observed for all the samples, while others are partially missing in a random pattern.
Classify. We included three deep learning algorithms for supervised classification tasks: Retrieval-AuGmented dynamic Prompt Tuning (RAGPT)29, Missing Modalities in Multimodal healthcare(M3Care)34, and MUtual-conSistEnt graph contrastive learning (MUSE)35. Notably, MUSE also supports training with missing labels. These methods are based on transformer architectures and leverage flexible pretrained models (such as Vision-and-Language Transformer, ViLT36), enabling them to classify datasets that include image, language, time-series, and tabular modalities.
Cluster. Clustering involves grouping samples into distinct groups. As shown in Fig. 1c, clustering is the most common task in MML with incomplete data, which is why this module offers the largest number of methods. It includes the following algorithms: Doubly Aligned Incomplete Multi-view Clustering (DAIMC)37, Efficient and Effective Incomplete Multi-view Clustering (EEIMVC)38, Incomplete Multiview Spectral Clustering With Adaptive Graph Learning (IMSCAGL)39, Integrate Any Omics (IntegrAO)40, Self-representation Subspace Clustering for Incomplete Multi-view Data (IMSR)41, Late Fusion Incomplete Multi-View Clustering (LFIMVC)42, Multiple Kernel k-Means with Incomplete Kernels (MKKMIK)43, Multi Omic clustering by Non-Exhaustive Types (MONET)44, Multi-Reconstruction Graph Convolutional Network (MRGCN)45, NEighborhood based Multi-Omics clustering (NEMO)46, Online multi-view clustering (OMVC)47, One-Pass Incomplete Multi-View Clustering (OPIMC)48, One-Stage Incomplete Multi-view Clustering via Late Fusion (OSLFIMVC)49, Projective Incomplete Multi-View Clustering (PIMVC)50, Scalable Incomplete Multiview Clustering with Adaptive Data Completion (SIMCADC)51 and Subtyping Tool for Multi-Omic data (SUMO)52. These methods utilize various approaches, including deep learning, matrix factorization, kernel methods, and graph learning, offering a comprehensive toolkit for clustering block-wise IMMD.
Explore. A diverse set of tools to explore a multi-modal dataset, such as reporting the number of complete and incomplete samples.
Decomposition. Reducing the number of features is often essential for preprocessing or exploring datasets. This module facilitates feature extraction by transforming the original feature space into a more compact representation. It includes the following algorithms: Data Fusion by Matrix Factorization (DFMF)53, Joint Non-negative Matrix Factorization Algorithms (jNMF)54, and Multi-Omics Factor Analysis (MOFA)55. These algorithms can handle both block- and feature-wise missing data. Once a model is fitted, any unseen data can be transformed using the learned parameters, making the extracted features available for downstream tasks such as supervised learning.
Impute. A module designed for filling missing data, which can be particularly useful when using external methods that are unable to handle missing values directly. Approaches using DFMF, MOFA and jNMF, mentioned earlier, are available for this task (see Supplementary Methods for details).
Feature selection. This module enables the identification of key features in IMMD. It offers a jNMF-based method that identifies the top features by analyzing their contributions to the underlying structure of the data (see Supplementary Methods).
Load. Classes in this module load and prepare datasets for deep learning algorithms.
Preprocessing. Utility functions and classes for processing and preparing IMMD for downstream tasks. It includes methods for selecting complete samples, dropping specific modalities, converting uni-modal transformations into multi-modal ones, among others.
Retrieve. We included a Multi-Channel Retriever29 designed to operate in vision-language scenarios with both full and missing modalities, identifying similar instances using pretrained models like CLIP (Contrastive Language Image Pre-training56). Additionally, it can generate retrieval-augmented prompts to support complementary tasks.
Statistics. Functions for quantifying multi-modal data statistics, such as the redundancy, uniqueness, and synergy of the modalities57.
Utils. Utilities for data manipulation, such as input validation.
Visualize. Functions to visually explore a multi-modal dataset.
We note that iMML does not automatically extend arbitrary external algorithms to handle missing modalities if they are not inherently designed for this purpose. However, iMML is a continuously evolving tool, and if future work proposes such an approach, we will add it to the framework.
Compatibility with other packages
iMML was designed to be compatible with widely used machine learning and data analysis tools, such as pandas58, NumPy59, Scikit-learn60, PyTorch61, and Lightning AI62, hence allowing researchers to apply machine learning models with minimal programming effort (Supplementary Fig. 9). This consistency allows iMML to be easily integrated into pipelines with other libraries that follow this style, offering high versatility and usability.
The input format that iMML expects is standardized, facilitating cross-comparisons of multiple methods with a unified interface (Supplementary Fig. 10).
Code quality
We have made a comprehensive effort to minimize external dependencies, ensuring that only well-maintained libraries are included. Continuous integration ensures compatibility with previous versions, and unit tests currently provide over 97% code coverage on Linux, Mac, and Windows platforms. Additionally, comprehensive documentation has been provided to make it easily usable.
We have introduced a random_state parameter across all algorithms in iMML, following the scikit-learn approach, which allows to obtain fully reproducible results. This feature was not present in most of the original implementations; thus, iMML ensures experimental reproducibility, which is a fundamental safeguard against the ongoing reproducibility crisis in science63.
Implementation analysis
To assess each translation, we tested five datasets (BBCSport64, BDGP65, BUAA66, Nutrimouse67 and sensIT30068) commonly used for assessing MML algorithms with incomplete data and ran each algorithm 50 times with varying random incomplete sample rates (Supplementary Methods 2.2). These experiments focused strictly on engine performance; thus, no preprocessing steps were applied, which could have significantly enhanced the overall results. We reported paired statistical comparisons between the original and translated implementations using the Nemenyi-Wilcoxon-Wilcox test to ensure that the new Python implementation does not deviate statistically from the original MATLAB/R code.
For the scalability analysis, we used the NUSWIDE dataset. This is a widely used real-world web image dataset with 5 modalities and 30,000 samples, in which each image is represented by five types of low-level features, i.e., color histogram, color correlogram, edge direction histogram, wavelet texture, and block-wise color moments69. We evaluated runtime and memory usage across 10 repeated runs and show how the performance changes with increasing sample size (300, 1500, 3000, 15,000 and 30,000), batch size (256, 512, 1024, 2048, 4096), hidden dimensions (32, 64, 128, 256, 512), number of epochs (1, 5, 10, 25, 50), and modality complexity (from 2 to 5 modalities).
Exploring an incomplete multi-omics dataset
TCGA pancreatic cancer data was downloaded from cBioPortal70 using pyBioPortal71. Partial information decomposition was executed with default hyperparameters, and missing values were replaced by the average of the feature57. Survival data was binarized at 18 months (selected between the median, 16, and the mean, 21). Censored patients alive earlier than 18 months were excluded. The method was applied 25 times with different initialization seeds to improve robustness, and the results were averaged. The best modality combination was selected based on total information, i.e., the combination providing the most information for survival prediction.
Classification and retrieval of an incomplete image-language dataset
RAGPT and M3Care were evaluated using a stratified five-fold cross-validation under the four missingness patterns and increasing incompleteness rates of 0%, 20%, 60%, and 80%. Because this analysis substantially increased computational requirements (approximately 1 week on an NVIDIA RTX 2080 Ti GPU), we used a binary subset of the Food101 dataset comprising the two most frequent classes (1424 samples)28. A linear model was fitted for each model to assess how performance drops when increasing missing data. The models were trained using their default hyperparameters and for 10 epochs with a batch size of 32 and 10 neighbors.
Clustering incomplete multi-modal data in text analysis
We simulated the four block-wise amputation patterns using iMML. For all methods, data were normalized prior to clustering. In each baseline, missing blocks were replaced with the feature-wise mean. The number of clusters was set to the actual number of classes in the dataset.
Imputation of block- and feature-wise incomplete multi-modal data in computer vision
We simulated the four block-wise amputation patterns using iMML, as well as random feature-wise missing data, and applied imputation. Both for MOFA and DFMF, data were z-score scaled before imputation, then reconverted to the original scale by applying an inverse transformation. The baseline imputation, which replaced missing values with the feature-wise mean, did not include this scaling step. The imputed multi-modal data from each method were then concatenated, z-score scaled, and used as inputs for logistic regression to predict their labels and compare with the ground truth (with balanced classes).
Feature extraction and feature selection on incomplete multi-modal data in biomedicine
We simulated various block- and feature-wise missing data patterns using iMML and applied the jNMF algorithm for feature extraction. For feature selection, the most influential feature from each component was chosen. To provide comparative benchmarks, we included baselines using randomly selected features and all available features. The outputs from these methods were then used as inputs for a support vector machine to predict genetic type (with balanced classes). As the feature selection process does not replace missing values, an imputation step was applied prior to classification. Relative modality importance was computed by summing the contributions of all features per modality, and then dividing them by the total.
Supplementary information
Source data
Author contributions
Conceptualization: A.L. Data curation: A.L. Formal analysis: A.L. Funding acquisition: T.A. Investigation: A.L. and T.D. Methodology: A.L., J.Z., and T.A. Software: A.L. and T.D. Supervision: J.Z. and T.A. Visualization: A.L. Writing—original draft: A.L. Writing—review and editing: A.L., J.Z., and T.A.
Peer review
Peer review information
Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.
Funding
This project has received funding from the European Union’s Horizon 2020 research and innovation program under grant agreement No 101016851, project PANCAIM, and from the Norwegian Cancer Society (grants 216104 and 273810), South-Eastern Norway Regional Health Authority (grants 2020026 and 2023105), Radium Hospital Foundation, the Finnish Cancer Foundation, and the Research Council of Finland (grants 340141 and 345803), under the frames of ERA PerMed (CLL-CLUE, grant 344698) and EP PerMed (CLL-OUTCOME, grant 367855).
Data availability
All datasets used in this study were publicly available: the BBCSport dataset is available at http://mlg.ucd.ie/datasets/segment.html, the BDGP dataset is available at http://ranger.uta.edu/h̃eng/Drosophila/data/, the BUAA dataset is available at https://github.com/hdzhao/IMG/tree/master/data, the Nutrimouse dataset is available at https://github.com/mvlearn/mvlearn/tree/main/mvlearn/datasets/nutrimouse, the sensIT300 dataset is available at https://github.com/Liuzhenjiao123/multiview-data-sets/blob/master/sensIT300.mat, the NUSWIDE dataset is available at https://drive.google.com/drive/folders/1O3YmthAZGiq1ZPSdE74R7Nwos2PmnHH, and the Statlog dataset is available at https://github.com/Liuzhenjiao123/multiview-data-sets/tree/master. TCGA data was downloaded from cBioPortal (https://www.cbioportal.org/) using pyBioPortal. The Food101 dataset is available at HuggingFace (https://huggingface.co/datasets/visual-layer/food101-vl-enriched). Source data are provided with this paper.
Code availability
iMML is provided as an open-source Python package at https://pypi.org/project/imml/, with public code https://github.com/ocbe-uio/imml and user instructions available at https://imml.readthedocs.io/. Code to perform the analyses in the manuscript and reproduce all the figures is available at https://zenodo.org/records/2048210572.
Competing interests
T.A. has received unrelated research funding from Mobius Biotechnology GmbH. The other authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Alberto López, Email: a.l.sanchez@medisin.uio.no.
Tero Aittokallio, Email: t.a.aittokallio@medisin.uio.no.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s41467-026-77212-w.
References
- 1.Huang, Y. et al. What makes multi-modal learning better than single (provably). In Proc. Advances in Neural Information Processing Systems Vol. 34 (eds Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P. & Vaughan, J. W.) 10944–10956 (Curran Associates, Inc., 2021).
- 2.Baltrusaitis, T., Ahuja, C. & Morency, L.-P. Multimodal machine learning: a survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell.41, 423–443 10.1109/TPAMI.2018.2798607 (2019). [DOI] [PubMed] [Google Scholar]
- 3.Girdhar, R. et al. ImageBind: one embedding space to bind them all. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 15180–15190 (IEEE, 2023).
- 4.Pei, X., Zuo, K., Li, Y. & Pang, Z. A review of the application of multi-modal deep learning in medicine: Bibliometrics and future directions. Int. J. Comput. Intell. Syst.16, 10.1007/s44196-023-00225-6 (2023). [DOI]
- 5.Xiang, C. et al. Multi-sensor fusion and cooperative perception for autonomous driving: a review. IEEE Intell. Transp. Syst. Mag.15, 36–58 (2023).
- 6.Fei, H. et al. From multimodal LLM to human-level AI: Modality, instruction, reasoning, efficiency and beyond. In Proc. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries (eds Klinger, R., Okazaki, N., Calzolari, N. & Kan, M.-Y.) 1–8 (ELRA and ICCL, 2024).
- 7.Wu, R., Wang, H., Chen, H.-T. & Carneiro, G. Deep multimodal learning with missing modality: A survey. Trans. Mach. Learn. Res. (2026).
- 8.Briere, G., Darbo, E., Thebault, P. & Uricaru, R. Consensus clustering applied to multi-omics disease subtyping. BMC Bioinform.22, 10.1186/s12859-021-04279-1 (2021). [DOI] [PMC free article] [PubMed]
- 9.Li, L., Zhang, J., Wang, Y. & Ran, B. Missing value imputation for traffic-related time series data based on a multi-view learning method. IEEE Trans. Intell. Transp. Syst.20, 2933–2943 10.1109/TITS.2018.2869768 (2019). [DOI] [Google Scholar]
- 10.Cohen, I., Cozman, F., Sebe, N., Cirelo, M. & Huang, T. Semisupervised learning of classifiers: theory, algorithms, and their application to human-computer interaction. IEEE Trans. Pattern Anal. Mach. Intell.26, 1553–1566 10.1109/TPAMI.2004.127 (2004). [DOI] [PubMed] [Google Scholar]
- 11.Tang, J., Yi, Q., Fu, S. & Tian, Y. Incomplete multi-view learning: review, analysis, and prospects. Appl. Soft Comput.153, 111278 10.1016/j.asoc.2024.111278 (2024). [DOI] [Google Scholar]
- 12.Zhao, J., Xijiong, X., Xu, X. & Sun, S. Multi-view learning overview: recent progress and new challenges. Inf. Fusion38, 10.1016/j.inffus.2017.02.007 (2017). [DOI]
- 13.Yan, X., Hu, S., Mao, Y., Ye, Y. & Yu, H. Deep multi-view learning methods: a review. Neurocomputing448, 106–129 (2021). [Google Scholar]
- 14.Kline, A. et al. Multimodal machine learning in precision health: a scoping review. npj Digit. Med.5, 171 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Khojaste-Sarakhsi, M., Haghighi, S. S., Ghomi, S. F. & Marchiori, E. Deep learning for alzheimer’s disease diagnosis: a survey. Artif. Intell. Med.130, 102332 10.1016/j.artmed.2022.102332 (2022). [DOI] [PubMed] [Google Scholar]
- 16.Steyaert, S. et al. Multimodal data fusion for cancer biomarker discovery with deep learning. Nat. Mach. Intell.5, 351–362 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Fang, U. et al. A comprehensive survey on multi-view clustering. IEEE Trans. Knowl. Data Eng.35, 12350–12368 10.1109/TKDE.2023.3270311 (2023). [DOI] [Google Scholar]
- 18.Huang, D., Wang, C.-D. & Lai, J.-H. Fast multi-view clustering via ensembles: towards scalability, superiority, and simplicity. IEEE Trans. Knowl. Data Eng.35, 11388–11402 10.1109/TKDE.2023.3236698 (2023). [DOI] [Google Scholar]
- 19.Chen, M.-S., Wang, C.-D. & Lai, J.-H. Low-rank tensor based proximity learning for multi-view clustering. IEEE Trans. Knowl. Data Eng.35, 5076–5090 10.1109/TKDE.2022.3151861 (2023). [DOI] [Google Scholar]
- 20.Tang, C. et al. Unified one-step multi-view spectral clustering. IEEE Trans. Knowl. Data Eng.35, 6449–6460 10.1109/TKDE.2022.3172687 (2023). [DOI] [Google Scholar]
- 21.Le, L. P., Nguyen, T., Riegler, M. A., Halvorsen, P. & Nguyen, B. T. Multimodal missing data in healthcare: a comprehensive review and future directions. Comput. Sci. Rev.56, 100720 10.1016/j.cosrev.2024.100720 (2025). [DOI] [Google Scholar]
- 22.Osipov, A. et al. The molecular twin artificial-intelligence platform integrates multi-omic data to predict outcomes for pancreatic adenocarcinoma patients. Nat. Cancer5, 1–16 10.1038/s43018-023-00697-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhou, Y., Bouadjenek, M. R. & Aryal, S. Missing data imputation: Do advanced ml/dl techniques outperform traditional approaches? In Proc. Machine Learning and Knowledge Discovery in Databases. Applied Data Science Track (eds Bifet, A., Krilavičius, T., Miliou, I. & Nowaczyk, S.) 100–115 (Springer Nature Switzerland, 2024).
- 24.Raphael, B. J. et al. Integrated genomic characterization of pancreatic ductal adenocarcinoma. Cancer Cell32, 185–203 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Tong, Y. et al. Proteogenomic insights into the biology and treatment of pancreatic ductal adenocarcinoma. J. Hematol. Oncol.15, 168 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Hyeon, D. Y. et al. Proteogenomic landscape of human pancreatic ductal adenocarcinoma in an asian population reveals tumor cell-enriched and immune-rich subtypes. Nat. Cancer4, 290–307 (2023). [DOI] [PubMed] [Google Scholar]
- 27.Miao, B., Lih, T.-S. M., Hu, Y. & Zhang, H. Signatures of pancreatic ductal adenocarcinoma uncovered by integrative multi-omics analysis. Cancers18, 687 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Bossard, L., Guillaumin, M. & Van Gool, L. Food-101 – mining discriminative components with random forests. In Proc. European Conference on Computer Vision 446–461 (Springer, 2014).
- 29.Lang, J., Cheng, Z., Zhong, T. & Zhou, F. Retrieval-augmented dynamic prompt tuning for incomplete multimodal learning. In Proc. 39th AAAI Conf. Artif. Intell. 2011 (AAAI Press, 2025).
- 30.Ren, L. Bbcsport, 10.21227/re2y-ph59 (2024). [DOI]
- 31.Statlog (Image Segmentation). UCI Machine Learning Repository. 10.24432/C5P01G (1990). [DOI]
- 32.Martin, P. G. P. et al. Novel aspects of pparα-mediated regulation of lipid and xenobiotic metabolism revealed through a nutrigenomic study. Hepatology45, 767–777 (2007). [DOI] [PubMed]
- 33.van Breugel, B. & van der Schaar, M. Position: Why tabular foundation models should be a research priority. In Proc. 41st Int. Conf. Mach. Learn. vol. 235, 48976–48993 (PMLR, 2024).
- 34.Zhang, C. et al. M3Care: Learning with missing modalities in multimodal healthcare data. In Proc. 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2418–2428 (ACM, 2022).
- 35.Wu, Z. et al. Multimodal patient representation learning with missing modalities and labels. In Proc. The Twelfth International Conference on Learning Representations. (ICLR, 2024).
- 36.Kim, W., Son, B. & Kim, I. ViLT: vision-and-language transformer without convolution or region supervision. In Proc. International Conference on Machine Learning 5583–5594 (PMLR, 2021).
- 37.Hu, M. & Chen, S. Doubly aligned incomplete multi-view clustering. In Proc. Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18 2262–2268 10.24963/ijcai.2018/313 (International Joint Conferences on Artificial Intelligence Organization, 2018). [DOI]
- 38.Liu, X. et al. Efficient and effective regularized incomplete multi-view clustering. IEEE Trans. Pattern Anal. Mach. Intell.43, 2634–2646 10.1109/TPAMI.2020.2974828 (2021). [DOI] [PubMed] [Google Scholar]
- 39.Wen, J., Xu, Y. & Liu, H. Incomplete multiview spectral clustering with adaptive graph learning. IEEE Trans. Cybern.50, 1418–1429 10.1109/TCYB.2018.2884715 (2020). [DOI] [PubMed] [Google Scholar]
- 40.Ma, S. et al. Moving towards genome-wide data integration for patient stratification with integrate any omics. Nat. Mach. Intell.7, 29–42 (2025). [Google Scholar]
- 41.Liu, J. et al. Self-representation subspace clustering for incomplete multi-view data. In Proc. 29th ACM International Conference on Multimedia, MM ’21 2726–2734 10.1145/3474085.3475379 (Association for Computing Machinery, 2021). [DOI]
- 42.Liu, X. et al. Late fusion incomplete multi-view clustering. IEEE Trans. Pattern Anal. Mach. Intell.41, 2410–2423 10.1109/TPAMI.2018.2879108 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Liu, X. et al. Multiple kernel kk-means with incomplete kernels. IEEE Trans. Pattern Anal. Mach. Intell.42, 1191–1204 10.1109/TPAMI.2019.2892416 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Rappoport, N., Safra, R. & Shamir, R. MONET: multi-omic module discovery by omic selection. PLOS Comput. Biol.16, e1008182 10.1371/journal.pcbi.1008182 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Yang, B., Yang, Y., Wang, M. & Su, X. MRGCN: cancer subtyping with multi-reconstruction graph convolutional network using full and partial multi-omics dataset. Bioinformatics39, 10.1093/bioinformatics/btad353 (2023). [DOI] [PMC free article] [PubMed]
- 46.Rappoport, N. & Shamir, R. NEMO: cancer subtyping by integration of partial multi-omic data. Bioinformatics35, 3348–3356 10.1093/bioinformatics/btz058 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Shao, W., He, L., Lu, C.-t. & Yu, P. S. Online multi-view clustering with incomplete views. In Proc. 2016 IEEE International Conference on Big Data (Big Data) 1012–1017 10.1109/BigData.2016.7840701 (2016). [DOI]
- 48.Hu, M. & Chen, S. One-pass incomplete multi-view clustering. In Proc. Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’19/IAAI’19/EAAI’19 10.1609/aaai.v33i01.33013838 (AAAI Press, 2019). [DOI]
- 49.Zhang, Y. et al. One-stage incomplete multi-view clustering via late fusion. In Proc. 29th ACM International Conference on Multimedia, MM ’21 2717–2725 10.1145/3474085.3475204 (Association for Computing Machinery, 2021). [DOI]
- 50.Deng, S. et al. Projective incomplete multi-view clustering. IEEE Trans. Neural Netw. Learn. Syst.35, 10539–10551 10.1109/TNNLS.2023.3242473 (2024). [DOI] [PubMed] [Google Scholar]
- 51.He, W. -j, Zhang, Z. & Wei, Y. Scalable incomplete multi-view clustering with adaptive data completion. Inf. Sci.649, 119562 10.1016/j.ins.2023.119562 (2023). [DOI] [Google Scholar]
- 52.Sienkiewicz, K. et al. Detecting molecular subtypes from multi-omics datasets using sumo. Cell Rep. Methods2, 100152 10.1016/j.crmeth.2021.100152 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Zitnik, M. & Zupan, B. Data fusion by matrix factorization. IEEE Trans. Pattern Anal. Mach. Intell.37, 41–53 10.1109/TPAMI.2014.2343973 (2015). [DOI] [PubMed] [Google Scholar]
- 54.Tsuyuzaki, K. & Nikaido, I. nnTensor: an R package for non-negative matrix/tensor decomposition. J. Open Source Softw.8, 5015 10.21105/joss.05015 (2023). [DOI] [Google Scholar]
- 55.Argelaguet, R. et al. Multi–omics factor analysis—a framework for unsupervised integration of multi–omics data sets. Mol. Syst. Biol.14, e8124 10.15252/msb.20178124 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Radford, A. et al. Learning transferable visual models from natural language supervision. In Proc. International Conference on Machine Learning 8748–8763 (PMLR, 2021).
- 57.Liang, P. P. et al. Quantifying & modeling multimodal interactions: an information decomposition framework. In Proc. Advances in Neural Information Processing Systems. Vol. 36, 27351–27393 (Curran Associates, 2023).
- 58.pandas development team, T. pandas-dev/pandas: Pandas, 10.5281/zenodo.3509134 (2020). [DOI]
- 59.Harris, C. R. et al. Array programming with numpy. Nature585, 357–362 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
- 61.Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems Vol. 32 (Curran Associates, 2019).
- 62.Falcon, W. & The PyTorch Lightning team. PyTorch Lightning. Zenodo 10.5281/zenodo.3828935 (2019). [DOI]
- 63.Cobey, K. D. et al. Biomedical researchers’ perspectives on the reproducibility of research. PLoS Biol.22, e3002870 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Greene, D. & Cunningham, P. Practical solutions to the problem of diagonal dominance in kernel document clustering. In Proc. 23rd International Conference on Machine Learning 377–384 (PMLR, 2006).
- 65.Cai, X., Wang, H., Huang, H. & Ding, C. Joint stage recognition and anatomical annotation of Drosophila gene expression patterns. Bioinformatics28, i16–i24 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Huang, D., Sun, J. & Wang, Y. The BUAA-VisNir Face Database Instructions. Technical Report No. IRIP-TR-12-FR-001, Vol. 3, 8 (School of Computer Science and Engineering, Beihang University, 2012).
- 67.Martin, P. G. et al. Novel aspects of pparα-mediated regulation of lipid and xenobiotic metabolism revealed through a nutrigenomic study. Hepatology45, 767–777 (2007). [DOI] [PubMed] [Google Scholar]
- 68.Duarte, M. F. & Hu, Y. H. Vehicle classification in distributed sensor networks. J. Parallel Distrib. Comput.64, 826–838 (2004). [Google Scholar]
- 69.Chua, T.-S. et al. NUS-WIDE: a real-world web image database from the National University of Singapore. In Proc. ACM International Conference on Image and Video Retrieval 1–9 (ACM, 2009).
- 70.Cerami, E. et al. The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer Discov.2, 401–404 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Valerio, M., Inno, A. & Gori, S. pybioportal: a Python package for simplifying cbioportal data access in cancer research. JAMIA Open8, ooae146 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.López, A. Multi-modal learning with incomplete data, 10.5281/zenodo.20482105 (2026). [DOI] [PubMed]
- 73.Wen, J. et al. A survey on incomplete multiview clustering. IEEE Trans. Syst., Man, Cybern. Syst.53, 1136–1149 10.1109/TSMC.2022.3192635 (2023). [DOI] [Google Scholar]
- 74.Liang, P. P. et al. Multizoo and multibench: a standardized toolkit for multimodal deep learning. J. Mach. Learn. Res.24, 1–7 (2023). [Google Scholar]
- 75.Benielli, D. et al. Toolbox for multimodal learn (scikit-multimodallearn). J. Mach. Learn. Res.23, 1–7 (2022). [Google Scholar]
- 76.Jarrett, D., Cebere, B. C., Liu, T., Curth, A. & van der Schaar, M. Hyperimpute: Generalized iterative imputation with automatic model selection. In Proc. International Conference on Machine Learning 9916–9937 (PMLR, 2022).
- 77.Schouten, R. M., Zamanzadeh, D. & Singh, P. pyampute: a Python library for data amputation, 10.25080/majora-212e5952-03e (2022). [DOI]
- 78.Schouten, R. M., Lugtig, P. & Vink, G. Generating missing values for simulation purposes: a multivariate amputation procedure. J. Stat. Comput. Simul.88, 2909–2930 (2018). [Google Scholar]
- 79.Perry, R. et al. mvlearn: multiview machine learning in Python. J. Mach. Learn. Res.22, 1–7 (2021). [Google Scholar]
- 80.Santos, M. S. et al. Generating synthetic missing data: a review by missing mechanism. IEEE Access7, 11651–11667 (2019). [Google Scholar]
- 81.van Buuren, S. & Groothuis-Oudshoorn, K. mice: multivariate imputation by chained equations in R. J. Stat. Softw.45, 1–67 10.18637/jss.v045.i03 (2011). [DOI] [Google Scholar]
- 82.Humanbased-AI. Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset. https://huggingface.co. Accessed 21 July 2026 (2026).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All datasets used in this study were publicly available: the BBCSport dataset is available at http://mlg.ucd.ie/datasets/segment.html, the BDGP dataset is available at http://ranger.uta.edu/h̃eng/Drosophila/data/, the BUAA dataset is available at https://github.com/hdzhao/IMG/tree/master/data, the Nutrimouse dataset is available at https://github.com/mvlearn/mvlearn/tree/main/mvlearn/datasets/nutrimouse, the sensIT300 dataset is available at https://github.com/Liuzhenjiao123/multiview-data-sets/blob/master/sensIT300.mat, the NUSWIDE dataset is available at https://drive.google.com/drive/folders/1O3YmthAZGiq1ZPSdE74R7Nwos2PmnHH, and the Statlog dataset is available at https://github.com/Liuzhenjiao123/multiview-data-sets/tree/master. TCGA data was downloaded from cBioPortal (https://www.cbioportal.org/) using pyBioPortal. The Food101 dataset is available at HuggingFace (https://huggingface.co/datasets/visual-layer/food101-vl-enriched). Source data are provided with this paper.
iMML is provided as an open-source Python package at https://pypi.org/project/imml/, with public code https://github.com/ocbe-uio/imml and user instructions available at https://imml.readthedocs.io/. Code to perform the analyses in the manuscript and reproduce all the figures is available at https://zenodo.org/records/2048210572.
