Skip to main content
Springer logoLink to Springer
. 2024 Dec 13;43(4):537–541. doi: 10.1007/s11604-024-01716-y

The critical need for an open medical imaging database in Japan: implications for global health and AI development

Daiju Ueda 1,✉, Shannon Walston 1, Hirotaka Takita 2, Yasuhito Mitsuyama 2, Yukio Miki 2
PMCID: PMC11953178  PMID: 39668276

Abstract

Japan leads OECD countries in medical imaging technology deployment but lacks open, large-scale medical imaging databases crucial for AI development. While Japan maintains extensive repositories, access restrictions limit their research utility, contrasting with open databases like the US Cancer Imaging Archive and UK Biobank. The 2018 Next Generation Medical Infrastructure Act attempted to address this through new data-sharing frameworks, but implementation has been limited by strict privacy regulations and institutional resistance. This data gap risks compromising AI system performance for Japanese patients and limits global medical AI advancement. The solution lies not in developing individual AI models, but in democratizing access to well-curated Japanese medical imaging data. By implementing privacy-preserving techniques and streamlining regulatory processes, Japan could enhance domestic healthcare outcomes while contributing to more robust global AI models, ultimately reclaiming its position as a leader in medical innovation.

Keywords: Open database, Artificial intelligence, Medical AI, AI fairness

Introduction

Japan has long been at the forefront of medical imaging technology, with its advanced diagnostic equipment being utilized in healthcare facilities worldwide. According to the Organization for Economic Co-operation and Development (OECD), Japan leads all member countries in the number of MRI and CT scanners per capita [1]. This technological advantage, however, stands in stark contrast to a significant shortcoming in Japan's contribution to global medical AI: the lack of open, large-scale medical imaging databases. The importance of such databases cannot be overstated in the current era of rapid AI advancement [2]. They serve as the foundation for developing and refining AI algorithms, which are increasingly becoming integral to medical diagnosis, treatment planning, and prognosis prediction [3]. The absence of Japanese data in these global AI training sets not only hinders Japan's potential to lead in medical AI but also has far-reaching implications for global healthcare advancements and AI fairness [4, 5].

Current state of medical imaging data in Japan

While Japan possesses vast repositories of high-quality medical imaging data, there is a notable absence of what could be termed "big data" in an open, accessible format. The Japan Medical Imaging Database (J-MID) represents Japan's impressive collection of medical imaging data [6]. However, despite this valuable resource, Japan currently lacks a truly open, large-scale medical imaging database. While J-MID maintains a vast collection, its access restrictions prevent it from fully serving the broader research community. This situation stands in contrast to other nations where unrestricted medical databases have become powerful catalysts for advancing AI innovation and research. For instance, the Cancer Imaging Archive (TCIA) in the United States provides researchers with access to large volumes of de-identified medical images of cancer patients [7]. Similarly, the UK Biobank offers researchers access to medical imaging data from hundreds of thousands of participants [8]. These open databases have significantly accelerated AI research in both their respective countries and globally.

The hesitancy toward open data sharing in Japan reflects cultural and institutional patterns. Unlike Western healthcare systems that have increasingly embraced data sharing as a driver of innovation, Japan's medical institutions traditionally view patient data as institutional property to be carefully guarded. This perspective stems from a historically paternalistic medical system and cultural emphasis on privacy protection. Moreover, Japan's post-war development model, which emphasized adapting existing technologies rather than pioneering new approaches, might have created an institutional culture that can be hesitant to lead in data-sharing initiatives.

The Japanese government has attempted to address this issue through the Next Generation Medical Infrastructure Act, enacted in 2018 and revised in 2024, and established new frameworks for medical data utilization [9, 10]. The 2024 revision introduced 'partially de-identified medical information' categories and guidelines for clinical trials and linkage analysis with the National Database of Health Insurance Claims. This legislation aimed to promote the utilization of medical data while protecting patient privacy. However, its implementation has faced several challenges. As of 2024, only a handful of medical institutions have been certified under this act to provide anonymized medical imaging [11]. The stringent requirements for data anonymization have proved to be a significant hurdle for many institutions. Moreover, there is a lack of public understanding about the importance of data sharing for medical research, leading to hesitancy in participation.

To address the difficulties of collecting and utilizing medical data, a working group was convened in 2022. This working group published the Guidelines on the Utilization of Medical Digital Data for AI Research and Development in 2024 [12]. Acknowledging the importance of collaborations between academic institutions and private companies, the new guidelines provide a comprehensive framework for healthcare institutions and private companies to collaborate on AI medical device development. These guidelines aim to maintain a balance between innovation and patient privacy rights.

Implications of the data gap

The scarcity of open Japanese medical imaging data has several critical implications for the overall robustness and generalizability of AI models. Recent studies have revealed troubling performance disparities in AI across different racial and gender groups [4, 13]. Seyyed-Kalantari et al. demonstrated that AI algorithms trained on chest radiographs showed lower performance for underrepresented patient populations [14]. With Japanese data largely absent from the training sets of global AI models, there is a real risk that these systems may underperform for Japanese patients, potentially compromising healthcare outcomes for Japanese citizens. Diversity in training data is crucial for developing AI systems that perform well across different populations [15]. By not contributing large-scale, open databases of Japanese medical imaging data, Japan is missing an opportunity to enhance the overall capabilities of global medical AI systems.

Recent advancements in Large Language Models (LLMs) [16–25] have further underscored the importance of this data gap. The integration of visual information and textual descriptions has become a cornerstone of cutting-edge medical AI applications [26, 27]. From automated diagnosis to personalized treatment planning and AI-assisted medical education, the ability to train models on paired image and language data is paramount. Japan's lack of such openly available datasets puts it at a significant disadvantage in this rapidly evolving field.

Challenges in creating open databases

Several factors contribute to the difficulty in establishing open medical imaging databases in Japan. Japan has stringent privacy laws and cultural norms that prioritize individual privacy [28, 29]. The Personal Information Protection Act, amended in 2020, sets strict guidelines for handling personal data, including medical information [30, 31]. While these protections are crucial, they can also create barriers to data sharing and open database creation. Many medical institutions in Japan are reluctant to share their data, primarily due to concerns over potential data misuse, such as breaches of privacy or unethical use, and the extensive effort required to prepare data for sharing. Data preparation involves several time-consuming tasks, such as anonymization, ensuring compliance with legal and ethical standards, and addressing compatibility issues across different systems. To overcome these challenges, it is necessary to establish appropriate incentives, such as financial support, recognition for contributing to research, or access to shared resources, to motivate institutions to participate in data sharing. This institutional inertia is a significant obstacle to creating comprehensive, open databases. Developing the necessary infrastructure for secure, efficient data sharing and management is a complex and costly undertaking. It requires not only technical expertise but also ongoing maintenance and updates to ensure data integrity and security [32]. While the Next Generation Medical Infrastructure Act was intended to facilitate medical data utilization, its implementation has revealed several regulatory challenges. The complex certification process for institutions and the strict anonymization requirements have limited its effectiveness.

Potential solutions and path forward

Addressing the challenges outlined above requires a multi-faceted approach. Implementing advanced anonymization and privacy-preserving techniques can help alleviate privacy concerns. Techniques, such as differential privacy, federated learning, defacing, and secure multi-party computation, can allow for data utilization while maintaining individual privacy [33]. Developing incentive structures for medical institutions to share their data could help overcome institutional resistance. This could include financial incentives, academic recognition, or priority access to developed AI tools [34]. Launching comprehensive public awareness campaigns is essential, particularly to clarify the personal information used, the value of utilizing this personal medical data in healthcare AI research, and the protective measures already in place for medical data sharing [35]. Such initiatives could help address common misconceptions about data misuse and build trust in medical AI development. Engaging in international collaborations and adopting global best practices in data sharing and AI development can help Japan align with the international standards while contributing its unique datasets to global research efforts [36]. Streamlining the regulatory process for data sharing and clarifying guidelines for anonymization and data usage could facilitate the creation of open databases while maintaining necessary protections.

All we need is the open dataset

The advancement of medical AI does not solely depend on developing sophisticated algorithms or individual models—it hinges critically on the availability of comprehensive, high-quality datasets. While creating specific AI solutions for medical challenges is valuable, the most transformative approach may lie in democratizing access to well-curated medical imaging data. When researchers worldwide face a particular medical challenge, the most effective catalyst for innovation often is not developing a single AI solution, but rather publishing diverse, representative datasets that encompass that challenge [7, 37]. The open dataset approach creates a multiplier effect, where hundreds of researchers can simultaneously work on solutions, leading to rapid iterations and improvements that would not be possible with siloed development.

The success of projects like ImageNet in computer vision and the Cancer Imaging Archive in medical imaging demonstrates how open datasets can accelerate scientific progress exponentially [7]. By making both data and associated code publicly available, we create an ecosystem where researchers can build upon each other's work, validate findings across different populations, and identify potential biases or limitations. This collaborative approach is particularly crucial in healthcare, where diverse patient populations and varying clinical contexts require robust solutions that can only emerge from extensive testing and refinement across multiple research groups [4].

Potential impact of open Japanese medical imaging databases

The establishment of open, large-scale medical imaging databases in Japan could have a far-reaching impact. By including Japanese data in AI training sets, models could be developed that perform more accurately and fairly for Japanese patients, potentially leading to better health outcomes [38]. The inclusion of diverse Japanese data in global datasets would contribute to the development of more robust and generalizable AI models, benefiting patients worldwide [39]. Fostering an environment conducive to medical AI research and development could attract international investment, stimulate domestic innovation, and potentially create new export opportunities in AI-driven healthcare solutions [40]. By taking a proactive role in medical data sharing and AI development, Japan could position itself as a leader in shaping international standards and ethical guidelines for AI in healthcare [41].

Conclusion

The establishment of open, large-scale medical imaging databases in Japan represents both a national imperative and a global necessity. It offers an opportunity for Japan to leverage its strengths in medical imaging technology, contribute meaningfully to the advancement of AI in healthcare, and ensure better health outcomes for its citizens and people worldwide. The path forward requires collaboration between government, healthcare institutions, and the tech industry. While challenges remain, particularly in the areas of privacy protection and institutional buy-in, the potential rewards—in terms of improved health outcomes, scientific advancement, and economic opportunity—make this an investment in the future that Japan cannot afford to overlook. By taking bold steps to open its vast medical imaging resources to the global research community, Japan can reclaim its position as a leader in medical innovation and play a pivotal role in shaping the future of global healthcare. The time for action is now, as the rapid pace of AI development in healthcare makes the need for diverse, comprehensive datasets more pressing than ever.

Acknowledgements

The English language in this manuscript was partially edited using Claude 3.5 Sonnet.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.OECD. Health at a Glance 2021 OECD indicators: OECD indicators. Paris: OECD Publishing; 2021. p. 2021. [Google Scholar]
  • 2.Ueda D, Shimazaki A, Miki Y. Technical and clinical overview of deep learning in radiology. Jpn J Radiol. 2019;37:15–33. [DOI] [PubMed] [Google Scholar]
  • 3.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25:44–56. [DOI] [PubMed] [Google Scholar]
  • 4.Ueda D, Kakinuma T, Fujita S, Kamagata K, Fushimi Y, Ito R, et al. Fairness of artificial intelligence in healthcare: review and recommendations. Jpn J Radiol. 2023;42:3–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Yoshiura T, Kiryu S. FAIR: a recipe for ensuring fairness in healthcare artificial intelligence. Jpn J Radiol. 2024;42:1–2. [DOI] [PubMed] [Google Scholar]
  • 6.Toshiaki A, Machitori A, Aoki S. Japan Radiological Society’s response to COVID-19 pneumonia and expectations for artificial intelligence in diagnostic radiology. Med Imaging Technol. 2021;39:3–7. [Google Scholar]
  • 7.Clark K, Vendt B, Smith K, Freymann J, Kirby J, Koppel P, et al. The cancer imaging archive (TCIA): maintaining and operating a public information repository. J Digit Imaging. 2013;26:1045–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Littlejohns TJ, Holliday J, Gibson LM, Garratt S, Oesingmann N, Alfaro-Almagro F, et al. The UK Biobank imaging enhancement of 100,000 participants: rationale, data collection, management and future directions. Nat Commun. 2020;11:2624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Matsuo R, Yamazaki T, Araki K. Development of a general statistical analytical system using nationally standardized medical information. J Med Syst. 2021;45:66. [DOI] [PubMed] [Google Scholar]
  • 10.Kumamaru H, Fukuma S, Matsui H, Kawasaki R, Tokumasu H, Takahashi A, et al. Principles for the use of large-scale medical databases to generate real-world evidence. Ann Clin Epidemiol. 2020;2:27–32. [Google Scholar]
  • 11.日本初、次世代医療基盤法に基づく医用画像データの提供開始. NTTデータ | Trusted Global Innovator. [cited 2024 Sep 24]. https://www.nttdata.com/global/ja/news/topics/2024/060701/
  • 12.厚生労働省. 医療デジタルデータのAI研究開発等への 利活用に係るガイドライン. [cited 2024 Nov 21]. https://www.mhlw.go.jp/content/001310044.pdf
  • 13.Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447–53. [DOI] [PubMed] [Google Scholar]
  • 14.Seyyed-Kalantari L, Zhang H, McDermott MBA, Chen IY, Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021;27:2176–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zou J, Schiebinger L. AI can be sexist and racist—it’s time to make it fair. Nature. 2018;559:324–6. [DOI] [PubMed] [Google Scholar]
  • 16.Ueda D, Mitsuyama Y, Takita H, Horiuchi D, Walston SL, Tatekawa H, et al. ChatGPT’s diagnostic performance from patient history and imaging findings on the diagnosis please quizzes. Radiology. 2023;308: e231040. [DOI] [PubMed] [Google Scholar]
  • 17.Horiuchi D, Tatekawa H, Shimono T, Walston SL, Takita H, Matsushita S, et al. Accuracy of ChatGPT generated diagnosis from patient’s medical history and imaging findings in neuroradiology cases. Neuroradiology. 2024;66:73–9. [DOI] [PubMed] [Google Scholar]
  • 18.Ueda D, Walston SL, Matsumoto T, Deguchi R, Tatekawa H, Miki Y. Evaluating GPT-4-based ChatGPT’s clinical potential on the NEJM quiz. BMC Digital Health. 2024;2:1–7. [Google Scholar]
  • 19.Oura T, Tatekawa H, Horiuchi D, Matsushita S, Takita H, Atsukawa N, et al. Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations. Jpn J Radiol. 2024. 10.1007/s11604-024-01633-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Mitsuyama Y, Tatekawa H, Takita H, Sasaki F, Tashiro A, Oue S, et al. Comparative analysis of GPT-4-based ChatGPT’s diagnostic performance with radiologists using real-world radiology reports of brain tumors. Eur Radiol. 2024. 10.1007/s00330-024-11032-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Nakaura T, Yoshida N, Kobayashi N, Shiraishi K, Nagayama Y, Uetani H, et al. Preliminary assessment of automated radiology report generation with generative pre-trained transformers: comparing results to radiologist-generated reports. Jpn J Radiol. 2024;42:190–200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Toyama Y, Harigai A, Abe M, Nagano M, Kawabata M, Seki Y, et al. Performance evaluation of ChatGPT, GPT-4, and Bard on the official board examination of the Japan Radiology Society. Jpn J Radiol. 2024;42:201–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Suzuki K, Yamada H, Yamazaki H, Honda G, Sakai S. Preliminary assessment of TNM classification performance for pancreatic cancer in Japanese radiology reports using GPT-4. Jpn J Radiol. 2024. 10.1007/s11604-024-01643-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Sonoda Y, Kurokawa R, Nakamura Y, Kanzawa J, Kurokawa M, Ohizumi Y, et al. Diagnostic performances of GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro in “Diagnosis Please” cases. Jpn J Radiol. 2024;42:1231–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hirano Y, Hanaoka S, Nakao T, Miki S, Kikuchi T, Nakamura Y, et al. GPT-4 Turbo with Vision fails to outperform text-only GPT-4 Turbo in the Japan Diagnostic Radiology Board Examination. Jpn J Radiol. 2024;42:918–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, et al. On the opportunities and risks of foundation models. arXiv [cs.LG]. 2021. http://arxiv.org/abs/2108.07258
  • 27.Horiuchi D, Tatekawa H, Oura T, Oue S, Walston SL, Takita H, et al. Comparing the diagnostic performance of GPT-4-based ChatGPT, GPT-4V-based ChatGPT, and radiologists in challenging neuroradiology cases. Clin Neuroradiol. 2024;34:779–87. [DOI] [PubMed] [Google Scholar]
  • 28.Rosen D. Private lives and public eyes: privacy in the United States and Japan. Fla J Int’l L. 1990;6:141. [Google Scholar]
  • 29.Cavoukian A, Chibba M. Privacy seals in the USA, Europe, Japan, Canada, India and Australia. In: Rodrigues R, Papakonstantinou V, editors. Privacy and data protection seals. The Hague: T.M.C. Asser Press; 2018. p. 59–82. [Google Scholar]
  • 30.令和2年 改正個人情報保護法について. [cited 2024 Sep 23]. https://www.ppc.go.jp/personalinfo/legal/kaiseihogohou/
  • 31.Naito Y, Aburatani H, Amano T, Baba E, Furukawa T, Hayashida T, et al. Clinical practice guidance for next-generation sequencing in cancer diagnosis and treatment (edition 2.1). Int J Clin Oncol. 2021;26:233–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Wilkinson MD, Dumontier M, Aalbersberg IJJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3: 160018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kaissis GA, Makowski MR, Rückert D, Braren RF. Secure, privacy-preserving and federated machine learning in medical imaging. Nat Mach Intell. 2020;2:305–11. [Google Scholar]
  • 34.Kostkova P, Brewer H, de Lusignan S, Fottrell E, Goldacre B, Hart G, et al. Who owns the data? Open data for healthcare. Front Public Health. 2016;4:7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Platt J, Kardia S. Public trust in health information sharing: implications for biobanking and electronic health record systems. J Pers Med. 2015;5:3–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Knoppers BM, Thorogood A, Chadwick R. The Human Genome Organisation: towards next-generation ethics. Genome Med. 2013;5:38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Qian B, Sheng B, Chen H, Wang X, Li T, Jin Y, et al. A competition for the diagnosis of myopic maculopathy by artificial intelligence algorithms. JAMA Ophthalmol. 2024;142:1006–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Park Y, Jackson GP, Foreman MA, Gruen D, Hu J, Das AK. Evaluating artificial intelligence in medicine: phases of clinical research. JAMIA Open. 2020;3:326–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH. Ensuring fairness in machine learning to advance health equity. Ann Intern Med. 2018;169:866–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Wolff J, Pauling J, Keck A, Baumbach J. The economic impact of artificial intelligence in health care: systematic review. J Med Internet Res. 2020;22: e16866. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Char DS, Shah NH, Magnus D. Implementing machine learning in health care—addressing ethical challenges. N Engl J Med. 2018;378:981–3. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Japanese Journal of Radiology are provided here courtesy of Springer

RESOURCES