Abstract
Objectives
Generative AI chatbots are revolutionizing health education by making complex information more accessible to the public. However, their use presents risks, including bias, hallucinations, ethical concerns, and misinformation, which are particularly critical in health contexts where incorrect guidance can have serious consequences. Ensuring safety, accuracy, and reliability is essential, especially in maternal and infant health.
Materials and Methods
We developed a multilingual chatbot, called Rosie, that employs a 3-stage AI pipeline, including a retriever, re-ranker, and generative model, to deliver efficient and relevant responses. To evaluate Rosie, we conducted a randomized controlled trial with pregnant and postpartum women (aged 14+ years, with infants under 6 months) from 49 US states (ClinicalTrials.gov ID NCT06053515). To analyze user interaction with Rosie, we examined 30 188 questions submitted by 197 users from October 2023 to April 2025. Six months after enrollment, a subset of Rosie participants were invited to complete a midpoint satisfaction survey assessing the chatbot’s response quality, clarity, usefulness, and feature engagement. Of 105 eligible participants, 84 completed the survey (80% completion rate) between November 2024 and June 2025.
Results
Users asked an average of 2.73 questions per day, with increased activity on weekdays and evenings, peaking at 10 p.m. The most highly rated topics included infant feeding, developmental milestones, symptoms (eg, fever), and hospital/birth preparations. Midpoint feedback was positive: 86% rated the answers as high quality, 91% found them useful, 95% found them easy to understand, and 84% expressed satisfaction.
Discussion
In the initial phase, Rosie users reported technical issues and less satisfactory responses. However, integrating a retrieval-augmented generation system, expanding Rosie’s knowledge base, and adding more interactive features led to sustained increases in positive ratings and consistently high user satisfaction. The Spanish-language expansion, enabled by a multilingual pipeline and advanced translation models, directly addressed pilot study feedback and further broadened Rosie’s accessibility. Rosie offers a broad approach, supporting users through pregnancy, childbirth, postpartum, and infant care during the first year.
Conclusion
Chatbots like Rosie have the potential to transform health information delivery by providing a scalable, personalized, and cost-efficient solution.
Keywords: chatbot, artificial intelligence, maternal health, health information, child health
Introduction
The pregnancy and postpartum periods are marked by significant physical and mental health changes. While these phases are critical for the well-being of both mother and newborn, many new mothers remain unaware of how to access timely, credible support resources after childbirth.1 Common concerns during this time include questions about nursing, newborn care, and managing the postpartum period. New mothers in underserved communities face disproportionate barriers to accessing high-quality healthcare and reliable postpartum information, contributing to poorer health outcomes and higher rates of infant mortality.2,3 These challenges are further exacerbated by the high prevalence of postpartum depression and anxiety, while language barriers also restrict access to essential therapeutic services.4,5 Although support services for new mothers exist, they are often limited by shortages of resources, personnel, funding, and the ability to provide personalized guidance. This gap often leads to the use of generic text messages that rarely meet the unique needs of each mother.
In contrast, health education chatbots offer a scalable, cost-effective solution with broad geographic reach. Chatbots have the potential to deliver personalized support, enable efficient data collection, and improve access to crucial information. For mothers and infants, such technological interventions can lead to improved health outcomes that extend across generations.6–8 Despite the recognized importance of promoting health literacy in underserved communities, a Maryland study on infant mortality identified a critical gap: the lack of health literacy initiatives that utilize innovative chatbot technology.9
Commercial chatbots, like ChatGPT, can provide timely and accessible answers to health questions, but their reliability remains uncertain. Key concerns with chatbots include bias, hallucinations (instances where false or fabricated information is presented as fact), limited knowledge, incorrect citations, and other ethical issues.10–12 Their outputs require validation and should be trained on diverse, reliable sources.13 A review of 128 studies found ChatGPT’s accuracy averages 73% but varies by specialty, performing better on simple tasks than complex clinical ones.14 While ChatGPT is promising for summarization and general medical information, large language models often lack reliability in complex health tasks and can spread misinformation or omit important caveats.15 Most health chatbots are English-based, though some multilingual chatbots like GyBot (Nepali/English)16–18 and MomConnect (11 South African languages)19 have shown promise for public health use but face limitations in scalability and advanced conversational skills. Other chatbots specifically designed for new mothers, like Kia-Chat, use a structured question-and-answer format where users select from predefined answer options, which can limit contextualized responses to personalized inquiries.20 As AI-powered technologies continue to evolve, numerous promising tools emerge. Nonetheless, there is an urgent need for the development of evidence-based, multilingual chatbots that can deliver accurate, real-time responses to questions related to pregnancy and child development.
Study aims
This paper describes the design and development of an interactive, Spanish and English chatbot (Rosie) to deliver evidence-based health information and prevent avoidable adverse health outcomes in mothers and newborns. This study evaluates whether Rosie, a bilingual AI-powered chatbot, is feasible and acceptable for delivering evidence-based maternal and infant health information to new mothers. We hypothesized that users would find Rosie’s answers useful and high-quality. We also hypothesized that positive feedback may be lower among less common health topics.
Methods
We developed a mobile application (Rosie) featuring a user-friendly interface and a core chatbot-based question-answering system. To ensure accurate and reliable baseline health information, we curated a comprehensive evidence-based library to create a corpus from which Rosie retrieves and synthesizes information to provide answers to users. We compiled over 60 sources from federal health agencies and professional health organizations such as the American Academy of Pediatrics (AAP),21 Centers for Disease Control and Prevention (CDC),22 children’s hospitals,23 MyHealthfinder,24 and KidsHealth.25 The corpus covered topics relevant to pregnancy and early parenting, including pregnancy complications, common health conditions, developmental milestones, nutrition and feeding, diapering, bathing, and available social support services.
Rosie’s question answering system
Rosie’s QA system was developed iteratively with user feedback and a human-in-the-loop process. Early versions used a 2-step retriever: one model selected the top 20 relevant passages from a knowledge bank, then another refined this to the top 10, returning the highest-ranked passage as the answer.26,27 While this ensured relevant, grounded responses, it struggled with generalization and coherence. To address this, we leveraged recent advances in question-answering28,29 by integrating Llama3.1,30 a state-of-the-art language model for question answering, which synthesizes contextualized answers from the top 10 passages. This approach merges the reliability of retrieval-based systems with the conversational fluency of generative models, making Rosie especially effective as a health chatbot. Rosie’s rating system was also updated to provide more detailed feedback. Initially, users could rate chatbot responses with a simple thumbs-up or thumbs-down. Later, this was replaced by a 5-star system that allowed users to evaluate responses based on relevance, accuracy, completeness, and citations, as well as provide optional written comments (Figure S1).
Evaluation of the reader model
To evaluate the reader model, annotators compared answers to 100 real user questions, assessing responses generated with and without the reader model. Annotators were all public health research assistants and faculty at the University of Maryland who received specific training on how to annotate. Annotators reviewed whether facts were supported by passages, attributions were correct, answers were topically relevant, and if they preferred the new answer. To enhance transparency and guard against hallucinations, generated answers included reference links and a visual similarity indicator. Sentence similarity to original passages was measured using TF-IDF and ROUGE-2 scores, with a yellow gradient underline highlighting higher similarity. Darker underlines indicate higher similarity, while lighter underlines suggest moderate similarity. This allows users to judge which parts of the answer are most reliably grounded in the source material.
Spanish language expansion
Given Spanish’s prominence in the U.S., a Spanish version of the Rosie app was created, featuring a language switch for users to ask and receive answers in Spanish. Limited high-quality Spanish health resources posed a challenge, so the team used neural machine translation (NMT) to convert the 600-million-character English corpus into Spanish.31–34 Using the OPUS-MT project and Marian NMT, English content was chunked and filtered before translation, which was run efficiently on GPUs.31,35 Manual evaluation showed Marian NMT was significantly more accurate than Google Translate36 (3 errors vs 20 in 100 samples), so the full corpus was translated with Marian NMT.
Rosie was enhanced for Spanish conversational support, enabling robust multilingual and dialectal dialogue. Rosie’s 3-stage Spanish QA pipeline includes (1) a multilingual retrieval model (mContriever)26 to find relevant passages, (2) a multilingual re-ranking model (Multilingual-E5-large)37 to refine results, and (3) a multilingual machine reading model (Llama3.1)30 to generate answers from the top passages. Figure 1 displays the timeline of Rosie’s mobile app development.
Figure 1.
Timeline of Rosie’s development from October 2023 to May 2025.
Rosie randomized trial
This study received ethics approval (#1556200) from the University of Maryland IRB, and the Rosie randomized clinical trial has been registered with ClinicalTrials.gov (NCT06053515). Eligible participants were mothers aged 14 and older, pregnant or parenting an infant under 6 months, and from racial or ethnic minority groups. Enrollment took place between October 2023 to December 2024, resulting in enrollment of 400 mothers across 49 states. Participants were randomized using a randomization table that specified the order in which participants were assigned to different treatment or control groups in the study. Half were randomized to receive the Rosie mobile app, while the other half (control group) received a children’s board book each month.
During enrollment, facilitators reviewed consent forms, compensation, and study duration. Participants were given the opportunity to ask questions to ensure full understanding before signing the form. Parental consent or presence was not required for minors per the Maryland Annotated Code § 20-102, which allows minors to receive information related to sexual and reproductive health without parental notification or consent, except in the case of sterilization or abortion, which were not included in the RCT. Written consent was obtained through the completion of a web-based form, and a copy of the consent form was sent via email to each enrolled participant. Rosie users received $0.50 per question (up to $45), with the top 20% of active users receiving a child’s educational gift set. Control group participants (Book Club) received one book per month for 12 months. Participants completed self-reported surveys via Qualtrics, with demographic information collected at enrollment and user engagement feedback gathered at the 6-month midpoint. In November 2024, additional engagement items were added for a subset of Rosie users. Research coordinators used mail, email, text, and phone follow-up to boost survey completion.
Additional features
The Rosie mobile app, available in English and Spanish, supports maternal and child health with features such as a searchable FAQ, resource lists, informational videos, and a peer discussion forum moderated for safety. The forum fosters supportive interaction among users, especially new mothers. The app also offers push notifications for appointment reminders and daily health tips, and allows users to reset their estimated delivery date for personalized content.
Ethical considerations
Data privacy safeguards include encrypting participant data in transit (eg, from users’ mobile phones to our research database) and at rest (in storage, not in use). Only research staff could access participant data, and the data was not shared with any third parties. Misinformation management was addressed by investing significant time in curating Rosie’s health information knowledge bank, using vetted sources such as the Mayo Clinic and children’s hospitals, resulting in the inclusion of 400 000 documents. Rosie was prompted to use this vetted knowledge bank to respond to users’ questions. However, if Rosie utilized other resources, this would be indicated in the responses (yellow-highlighted text had a high probability of originating from Rosie’s vetted knowledge base, and alternatively, unhighlighted text had a high probability of originating from external sources). All statements had web reference links.
Data cleaning and analytic approach
To analyze user interaction with Rosie, we examined 30 188 questions submitted by 197 users from October 2023 to April 2025 (Note: 3 participants assigned to the Rosie group did not ask any questions). Questions containing terms related to babies or children were categorized as “baby-specified,” those with pregnancy or mother-related terms as “mother-specified,” questions with both as “both,” and those with neither as “other.” Additionally, after removing stop words and tokenizing the text, we conducted descriptive analyses to identify frequent topics and assess submission patterns by time of day. We also analyzed user ratings of Rosie’s responses to identify satisfaction trends. To explore feedback further, we applied topic modeling. Questions were preprocessed, then Uniform Manifold Approximation and Projection (UMAP)38 and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN)39 were used for dimensionality reduction and clustering. BERTopic40 generated interpretable topic clusters, which were manually reviewed for accuracy.
We also applied Latent Dirichlet Allocation (LDA),41 another topic modeling method, to all user ratings. The model generated word clusters for each topic, which the study team manually interpreted, labeled, and reviewed based on their overall meaning. For example, a cluster containing “milk,” “breast,” and “gas” was labeled as feeding related. Because LDA captures language patterns but not context, manual labeling and team discussion were necessary to ensure that the final themes accurately reflected maternal and infant domains (Table S1).
Additionally, 6 months after enrollment, a subset of Rosie participants were invited to complete a midpoint satisfaction survey assessing the chatbot’s response quality, clarity, usefulness, and feature engagement. Of 105 eligible participants, 84 completed the survey (80% completion rate) between November 2024 and June 2025.
Results
Rosie app structure and user feedback mechanism
Table 1 outlines the Rosie mobile application components, including its core functionalities and language support features. Figure 2 showcases the visual appearance of each Rosie app feature in English, while Figure S2 displays the visual appearance in Spanish.
Table 1.
Rosie mobile application components.
| Category | Component | Description |
|---|---|---|
| Application Pages | Rosie Chat Page | An interactive page where users can ask questions and receive answers from Rosie. |
| Daily Tips Page | Displays daily health tips and tips history for the past 7 days. | |
| FAQ Page | Static question-and-answer pairs are provided and organized by categories. | |
| Video Library Page | Offers educational videos on various topics, including sleeping and bathing. | |
| Discussion Board | Provides a space for users to connect, share experiences, and ask questions of one another. | |
| Profile Page | Allows users to manage personal settings and change language preferences. | |
| Functionalities | Push Notification | Reminds users to use the app. Moreover, Rosie tracks the baby’s age and reminds users of upcoming well-baby appointments (3-5 days after birth, 1 month, 2 months, 4 months, 6 months, 9 months, and 12 months). |
| Delivery Date Reset | Allow users to update the expected delivery date to personalize daily tips. | |
| Language Switch | Enables users to toggle between English and Spanish. | |
| Feedback System | Users can rate each response that Rosie provides, leaving feedback on what they like or dislike about the responses. | |
| Participation Reward System | Basic counts for participation (number of questions asked) | |
| Account Deletion | Users can submit deactivation requests for the Rosie app within the app. Our team reviews each request and contacts users. |
Figure 2.
English Rosie mobile application. The panel from left to right: (A) Login page where users can log in using email or Apple ID/Android ID. (B) Rosie chat page where users can ask Rosie a health question, and an example of Rosie’s response to a health question with web sources displayed as references and options for users to rate the response. (C) Daily tips page, (D) FAQ, (E) video library, and (F) discussion forum.
Evaluation of the reader model
Results from annotators who evaluated the reader model demonstrated a clear preference for the reader-enhanced answers:
90% were judged topically relevant.
91% had all facts grounded in the source passages.
88% had accurate attributions.
89% were preferred overall.
These results suggest improved contextual grounding, clarity, and user satisfaction compared to the earlier question-answer model.
User engagement with the Rosie app
Figure S3 shows that Rosie users averaged about 4 questions per day at enrollment, dropping to around 2 per day by days 30-90, then stabilizing. Usage spikes occurred around days 180, 265, and 320 and peaked to 5 questions on day 347, likely reflecting user exploration, normalization, and re-engagement (possibly due to app updates). The overall 7-day rolling average was 2.73 questions per day. Engagement was higher on weekdays and evenings and peaked on Tuesdays at 10 p.m. (Figures S4 and S5).
Midpoint user survey results
We analyzed the midpoint survey responses from November 2024 to June 2025. The sample was diverse in race, ethnicity, motherhood status, and income; most participants had at least a high school education, were insured, and the average age was 32 years (Table 2).
Table 2.
Rosie participant characteristics.
| Characteristic | Number of participants (n) | Percentage (%) |
|---|---|---|
| Parenting status | ||
| Pregnant | 33 | 39 |
| Parenting | 51 | 61 |
| Race/ethnicity | ||
| American Indian/Alaska Native | 6 | 7 |
| Asian | 18 | 21 |
| Black/African American | 33 | 39 |
| Hispanic/Latino | 19 | 22 |
| Multiracial | 8 | 10 |
| Ethnicity | ||
| Hispanic/Latino | 23 | 27 |
| Non-Hispanic/Latino | 61 | 73 |
| Education level | ||
| Less than High School | 2 | 2 |
| High School | 17 | 20 |
| Associate’s | 3 | 4 |
| Bachelor’s | 29 | 34 |
| Master’s | 25 | 30 |
| Doctorate/Professional | 8 | 10 |
| Household income | ||
| $0-30K | 19 | 23 |
| $30K-50K | 8 | 10 |
| $50K-75K | 14 | 17 |
| $75K-100K | 16 | 19 |
| $100K-200K | 18 | 21 |
| $200K+ | 7 | 8 |
| Prefer not to say | 2 | 2 |
| Insurance status | ||
| Insured | 79 | 94 |
| Uninsured | 5 | 6 |
| Age (years, Mean [SD]) | 32 (5) | |
Values are reported as n (%) unless otherwise specified.
SD = standard deviation. Of 105 eligible participants, 84 completed the 6-month Rosie mid-point survey (80% completion rate) between November 2024 and June 2025.
Users reported that Rosie’s answers were highly rated in quality, usefulness, ease of understanding, and satisfaction (Figure 3).
Figure 3.
Rosie users’ feedback at 6-month midpoint survey. A bar chart of users’ ratings of Rosie’s for the following: (A) answers to health questions were of high quality, (B) answers were easy to understand, (C) useful as a parent, and (D) satisfaction with answers.
User engagement (n = 80) was self-reported from “Never” to “Most of the time” (Figure S6). Most users clicked on answer source links and rated responses at least occasionally, with 22 and 40 users doing so “most of the time,” respectively. The most-used feature was tracking points for asking questions, with over 39% checking weekly. Other features like the video library, FAQ, and daily syllabus were used less frequently. Common technical issues reported at the 6-month midpoint by users (n = 80) included app crashes (35%), slow response times (37%), difficulties updating the app (35%), unexpected logouts (14%), and excessive storage space consumption (10%).
Analysis of questions submitted by users
Analysis of all 30 188 submitted user questions from October 2023 to April 2025 showed that 42% were fragments rather than full questions, indicating Rosie could interpret intent regardless of phrasing. The most common question types were “how” (22%), “what” (19%), “when” (11%), and “why” (5%), with “how” questions often seeking detailed, practical guidance. Categorization revealed 48% of questions were “baby-specified,” 42% were “other,” and only 10% were “mother-specified.”
User evaluation of the answers provided by the Rosie app
User feedback on Rosie’s responses (13 008 thumbs-up; 5831 thumbs-down) showed a widening gap favoring positive ratings (Figure 4). Statistical analysis confirmed a significant weekly increase in good feedback (Sen’s slope Q = +0.18 percentage points per week, non-parametric Mann-Kendall test P < .001), indicating steady improvements in Rosie’s performance and usability over time. A weighted least squares (WLS) model that adjusted for heteroscedasticity also confirmed the statistically significant upward trend.
Figure 4.
Rosie app satisfaction. Weekly percentages of ‘good’ (satisfactory) and ‘bad’ (unsatisfactory) ratings, spanning from October 2023 to April 2025, are represented by blue and pink, respectively.
Figure 5 shows that most user ratings were positive across the top 20 broad topic categories generated by BERTopic. Satisfactory feedback was highest for core maternal and infant health topics like feeding, child development, and common symptoms—topics well-covered in Rosie’s corpus. Unsatisfactory ratings were more common for narrow or less-represented topics requiring specialized knowledge, such as specific symptoms, travel, or herbal remedies. For example, food-related questions (mostly about baby feeding) received 73% positive ratings, while pain-related (mainly maternal) questions had lower positive ratings (61%) (Figures S7 and S8). Nonetheless, positive feedback improved significantly over time for pain-related questions (Mann-Kendall P < .001; Sen’s slope Q = +0.13 pp/week) and food-related questions (Mann-Kendall P < .001; Sen’s slope Q = +0.05 pp/week).
Figure 5.
Rosie’s response ratings. Users rated Rosie’s responses as “good” (satisfactory) or “bad” (unsatisfactory). The proportions of “good” and “bad” ratings are depicted in green and pink, respectively.
We also analyzed feedback from “super users,” whose submissions exceeded the upper IQR outlier threshold (Q3+1.5×IQR) to track changes over time (4% of trial Rosie users). Figure S9 shows that, for most super users, positive feedback increased especially after model updates while negative feedback declined, or the gap narrowed. These trends suggest user perceptions of Rosie generally improved with continued use and model enhancements.
Discussion
In the initial phase, Rosie users reported technical issues and less satisfactory responses.3 However, integrating a retrieval-augmented generation system, expanding Rosie’s knowledge base, and adding more interactive features led to sustained increases in positive ratings and consistently high user satisfaction. The Spanish-language expansion, enabled by a multilingual pipeline and advanced translation models, directly addressed pilot study feedback and further broadened Rosie’s accessibility. Unlike other chatbots focused solely on the perinatal period,42 Rosie offers a broader approach, supporting users through pregnancy, childbirth, postpartum, and infant care during the first year. Compared with these perinatal applications, which reported medians of only 6 logins over 12 months or single sessions lasting 13-20 minutes, Rosie demonstrated intensive and sustained use, accompanied by high satisfaction rates (84%-95%) and a steady improvement in positive ratings over time. Rosie offers flexible, evidence-based interactions, surpassing earlier rule-based or narrowly focused models.43 Features like underlined similarity to source documents and direct links increased answer transparency and trust. Midpoint surveys indicated high satisfaction with Rosie’s response quality, clarity, and usefulness. User interaction data revealed that most questions were submitted during evenings, highlighting the demand for health support outside typical clinic hours.44 The most common queries began with “how,” “what,” and “why,” showing users’ interest in actionable guidance and clear explanations. Topic modeling confirmed Rosie’s strongest performance on foundational maternal and infant health topics such as feeding, developmental milestones, labor, and infant care, reflecting the breadth and reliability of its training corpus.
Several health chatbots have been piloted internationally with promising results. In Australia, a chatbot designed to assess alcohol consumption and provide educational feedback using a factual response database received positive feedback in a trial with 17 users.45 In Switzerland, another chatbot proved useful in obesity treatment interventions.46 In the United States, a Massachusetts pilot study evaluated a chatbot that delivered breastfeeding information. The study found that users preferred being able to ask open-ended questions over selecting from preset choices, and they felt comfortable with the chatbot because its information was medically verified.47 Notably, users reported a preference for discussing sensitive issues such as mental health, breastfeeding difficulties, body image, and sexual health with a chatbot rather than a doctor, citing reduced fear of judgment and greater opportunities for dialogue compared to brief clinical visits. These findings highlight the potential of chatbots to offer a supportive, nonjudgmental environment for sensitive health topics.47 While our study did not explicitly compare user preferences for discussing sensitive issues with Rosie vs a doctor, we observed that participants frequently asked about topics like breastfeeding, mental health, anxiety, and postpartum depression indicating that Rosie is used for addressing a broad range of sensitive concerns.
Strengths of the approach
A key strength of the Rosie chatbot is its interdisciplinary development framework, which brings together expertise in maternal well-being, child health, and artificial intelligence. Rosie’s knowledge base is constructed from rigorously vetted clinical sources, including the CDC, NIH, Mayo Clinic, and AAP, making it more reliable than many commercial chatbots, which often rely on unvetted data.48 This foundation enables Rosie to deliver robust chatbot performance and provide enhanced user experiences. Rosie’s interactive features (video library, FAQ section, daily tips, answer rating system, and discussion board) support multiple learning modalities and foster greater user engagement. Continuous user feedback and participation facilitated iterative refinement of the model. Understanding users’ preferred app usage times and the topics that receive high or low ratings helped guide decisions on when to introduce new features or allocate additional support.
While maternal and reproductive health chatbots have shown promising results, they still face challenges that limit their long-term impact. For example, “Woebot Mom,” a digital program for pregnant and postpartum women, demonstrated measurable reductions in depressive symptoms, but its focus was limited to mental health, overlooking the broader and interconnected needs of mothers and infants.49 The GISSA Mother-Baby chatbot in Brazil and Parentbot in Singapore reported improvements in maternal satisfaction and parenting self-efficacy, yet these outcomes were observed over short evaluation periods and lacked clear strategies for integration into formal health systems, raising concerns about sustainability.50 Other tools, such as Dr Joy in South Korea and Gabby in the United States, were valued as educational resources for obstetric care and preconception risk reduction, but users often described them as “robotic” or insufficiently tailored to individual and cultural contexts.42 These existing solutions continue to struggle with generic responses, limited empathy, technical glitches, and the absence of continuous feedback or human-in-the-loop mechanisms for refining model performance. Although they improve access and provide judgment-free spaces, their fragmented design limits their ability to deliver lasting improvements in maternal and child health.42,50
Limitations and challenges
Despite its strengths, this study has several limitations. Topics with low user ratings suggest that Rosie may not fully address the broad spectrum of maternal and child health concerns, likely due to the limited size of its training dataset. Technical issues, such as long response times caused by using a more computationally intensive model and occasional app crashes, were also reported and may negatively impact user experience. Ongoing, iterative improvements to the app are necessary to optimize usability. The scarcity of Spanish-language health information posed additional challenges in developing a multilingual system. Rigorous evaluation of translation models was essential for building the Spanish corpus, and further linguistic validation is needed to ensure high accuracy. Currently, a pilot study with Spanish-speaking Rosie users is underway, which will provide valuable insights into the app’s effectiveness in Spanish.
Expanding the use of health chatbots like Rosie beyond the U.S. context introduces further complexities. Chatbots may lack support for local languages, dialects, or culturally appropriate phrasing. In addition, health concepts such as mental or reproductive health can carry stigma or be interpreted differently across cultures. Accurate language translation and careful tailoring of health information are crucial to ensure content is both understandable and culturally relevant for the intended user community. Moreover, access to smartphones, reliable internet, or electricity remains limited in some regions. The populations that could benefit most, such as rural communities, older adults, and individuals with low literacy, are often those least able to access chatbot technology. Health literacy barriers may further hinder effective use.
For future trials involving health information chatbots for parents, it is recommended to assess not only changes in health outcomes but also changes in health knowledge, parental self-efficacy, and health behaviors. Given that usage of digital health tools often declines sharply after initial engagement, researchers should plan data collection at both short- and long-term intervals (eg, at 1, 2, 3, 6, and 12 months). This time-sensitive approach is particularly important, as certain changes, such as fluctuations in mental health during pregnancy and postpartum, may be temporary but occur during critical periods for mothers and their children.
Conclusions
Traditional health information strategies have typically relied on one-way communication, often failing to tailor content to individual needs. Chatbots like Rosie offer a promising alternative by enabling bidirectional engagement and providing personalized health information. Our findings demonstrate the feasibility and utility of chatbot-based interventions in delivering real-time, accurate, and customized support to women during the gestational and postpartum periods. These stages, frequently overlooked by traditional healthcare systems, are critical for meeting the informational needs of new mothers particularly among racial and ethnic minorities, who are disproportionately affected by adverse maternal and infant health outcomes. Rosie’s design not only ensures the delivery of reliable information but also fosters active user participation and engagement. The platform’s interactive features enable the collection of valuable user feedback, which informs ongoing improvements and iterations of the tool. Future work should focus on expanding Rosie’s reach to broad geographic and cultural contexts, and on evaluating both short- and long-term impacts. Additionally, further development should enhance Rosie’s ability to provide personalized support tailored to users’ specific maternal health needs. Such efforts will be essential for improving health outcomes for women during pregnancy and the postpartum period.
Supplementary Material
Acknowledgment
We thank Vidur Jain for his help with this manuscript.
Contributor Information
Quynh C Nguyen, National Institute of Nursing Research (NINR), National Institutes of Health (NIH), Bethesda, MD 20892, United States.
Elizabeth M Norell, Department of Behavioral and Community Health, University of Maryland School of Public Health, College Park, MD 20742, United States.
Heran Mane, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Xiaohe Yue, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Neha Pundlik Srikanth, Department of Computer Science, UMIACS, University of Maryland, College Park, MD 20742, United States.
Carson J Peters, Department of Behavioral and Community Health, University of Maryland School of Public Health, College Park, MD 20742, United States.
Francia Ximena Marin Gutierrez, Department of Behavioral and Community Health, University of Maryland School of Public Health, College Park, MD 20742, United States.
Eesha Kurella, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Pankaj Dipankar, National Institute of Nursing Research (NINR), National Institutes of Health (NIH), Bethesda, MD 20892, United States.
Diego Salazar, National Institute of Nursing Research (NINR), National Institutes of Health (NIH), Bethesda, MD 20892, United States.
Penchala Sai Priya Mullaputi, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Amrutha Alibilli, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Adwaith Santhosh, Department of Data Science, College of Computer, Mathematical, and Natural Sciences, University of Maryland, College Park, MD 20742, United States.
Xin He, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Jordan Boyd-Graber, Department of Computer Science, UMIACS, University of Maryland, College Park, MD 20742, United States.
Thu T Nguyen, Department of Epidemiology and Biostatistics, University of Maryland School of Public Health, College Park, MD 20742, United States.
Author contributions
Quynh C. Nguyen (Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Writing—original draft, Writing—review & editing), Elizabeth M. Aparicio (Funding acquisition, Project administration, Resources, Supervision, Writing—original draft, Writing—review & editing), Heran Mane (Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing—original draft, Writing—review & editing), Xiaohe Yue (Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing—original draft, Writing—review & editing), Neha P. Srikanth (Conceptualization, Formal analysis, Investigation, Methodology, Software, Supervision, Validation, Writing—original draft, Writing—review & editing), Carson Peters (Data curation, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing—original draft, Writing—review & editing), Francia Ximena Marin Gutierrez (Data curation, Project administration, Resources, Supervision, Writing—original draft, Writing—review & editing), Eesha Kurella (Formal analysis, Investigation, Methodology, Validation, Writing—original draft, Writing—review & editing), Pankaj Dipankar (Formal analysis, Visualization, Writing—original draft, Writing—review & editing), Diego Salazar (Writing—original draft, Writing—review & editing), Penchala Sai Priya Mullaputi (Writing—original draft, Writing—review & editing), Amrutha Alibilli (Writing—original draft, Writing—review & editing), Adwaith Santhosh (Formal analysis, Methodology, Writing—review & editing), Xin He (Formal analysis, Funding acquisition, Investigation, Methodology, Supervision, Writing—original draft, Writing—review & editing), Jordan Boyd-Garber (Conceptualization, Funding acquisition, Methodology, Resources, Software, Supervision, Writing—review & editing), and Thu T. Nguyen (Funding acquisition, Project administration, Supervision, Writing—original draft, Writing—review & editing)
Supplementary material
Supplementary material is available at JAMIA Open online.
Funding
The work reported in this publication was supported by the National Institute on Minority Health and Health Disparities: R01MD016037 and R01MD015716. Additionally, this research was supported [in part] by the Intramural Research Program of the National Institutes of Health (NIH) (grant number ZIA NR000043; PI Nguyen). The contributions of the NIH author(s) are considered Works of the United States Government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Conflicts of interest
The authors declare that they have no conflicts of interest.
Data availability
The de-identified versions of data sets generated during and/or analyzed during this study are available from the corresponding author upon reasonable request.
References
- 1. MacArthur C, Winter H, Bick D, et al. Effects of redesigned community postnatal care on womens’ health 4 months after birth: a cluster randomised controlled trial. Lancet. 2002;359:378-385. [DOI] [PubMed] [Google Scholar]
- 2. Mane HY, Doig AC, Gutierrez FXM, et al. Practical guidance for the development of rosie, a health education question-and-answer chatbot for new mothers. J Public Health Manag Pract. 2023;29:663-670. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Nguyen QC, Aparicio EM, Jasczynski M, et al. Rosie, a health education question-and-answer chatbot for new mothers: randomized pilot study. JMIR Form Res. 2024;8:e51361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. AlMakinah R, Norcini-Pala A, Disney L, et al. Enhancing mental health support through human-AI collaboration: toward secure and empathetic AI-enabled chatbots. 2025 IEEE Conference on Artificial Intelligence (CAI). IEEE; 2025:196-202.
- 5. Batani J, Mbunge E, Leokana L. A deep learning-based chatbot to enhance maternal health education. 2024 Conference on Information Communications Technology and Society (ICTAS). IEEE; 2024:7-11.
- 6. Sowmya D, Debbata K, Vignesh GS, et al. Emerging role of healthcare chatbots in improving medical assistance. 2023 3rd Asian Conference on Innovation in Technology (ASIANCON). IEEE; 2023:1-6.
- 7. Dohare P, Johri S, Priya S, et al. Good fellow: a healthcare chatbot system. 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT). IEEE; 2023:1-6.
- 8. Goel R, Goswami RP, Totlani S, et al. Machine learning based healthcare chatbot. 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE). IEEE; 2022:188-192.
- 9. Renbarger KM, Abebe S, Place JM, et al. Perspectives of infant mortality from African American community members. Womens Health Rep (New Rochelle). 2023;4:423-430. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Shetty V, Gul A, Patel H, et al. Role of ChatGPT in antenatal care. World J Adv Res Rev. 2024;23:230-236. [Google Scholar]
- 11. Ayo-Ajibola O, Julien C, Lin ME, et al. Association of primary care access with health-related chatgpt use: a national cross-sectional survey. J Gen Intern Med. 2025. 10.1007/s11606-025-09406-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Sallam M. ChatGPT utility in healthcare education, research, and practice: Systematic review on the promising perspectives and valid concerns. Healthcare. 2023;11:887. 10.3390/healthcare11060887 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Özer Aslan İ, Aslan MT. Benchmarking AI chatbots for maternal lactation support: a cross-platform evaluation of quality, readability, and clinical accuracy. Healthcare. 2025;13:1756. 10.3390/healthcare13141756 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Beheshti M, Toubal IE, Alaboud K, et al. Evaluating the reliability of chatgpt for health-related questions: a systematic review. Informatics. 2025;12:9. 10.3390/informatics12010009 [DOI] [Google Scholar]
- 15. Wang L, Wan Z, Ni C, et al. Applications and concerns of ChatGPT and other conversational large language models in health care: systematic review. J Med Internet Res. 2024;26:e22769. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Kaneho AEA, Zrira N, Ouazzani-Touhami K, et al. Development of a bilingual healthcare chatbot for pregnant women: a comparative study of deep learning models with BiGRU optimization. Intell Based Med. 2025;12:100261. [Google Scholar]
- 17. Poudel S, Ghimire N, Subedi B, et al. Retrieval and generative approaches for a pregnancy chatbot in Nepali with stemmed and non-stemmed data: a comparative study. In: Proceedings of the International Conference on Technologies for Computer, Electrical, Electronics & Communication (ICT-CEEL 2023). Bhaktapur, Nepal; 2023. 10.48550/arXiv.2311.06898 [DOI]
- 18. Elwahsh S, Stern N, Singh A, et al. Linguistic diversity and mental well-being: co-designing custom ai chatbots with multilingual mothers. Proceedings of the 7th ACM Conference on Conversational User Interfaces. Association for Computing Machinery (ACM); 2025:1-17.
- 19. MomConnect. National Digital Maternal Health Program in South Africa. 2024. Accessed September 26, 2025. https://en.wikipedia.org/wiki/MomConnect
- 20. Vinarti RA, Sani NA, Amalia R, et al. KIA-CHAT: a QnA chatbot for postnatal and newborn care. JPK. 2024;12:97-101. 10.20473/jpk.V12.ISI1.2024.97-101 [DOI] [Google Scholar]
- 21. American Academy of Pediatrics PSoC, Holistic, and Integrative Medicine. Accessed November 29, 2006. www.aap.org/sections/chim
- 22. CDC. Centers for Disease Control and Prevention. Accessed January 25, 2025. http://www.cdc.gov/
- 23. NACHRI. All children need children’s hospitals. NACHRI; 2001. [Google Scholar]
- 24. ODPHP. Office of Disease Prevention and Health Promotion, U.S. DHHS. myHealthfinder Developer API; 2016. Accessed February 12, 2016. http://healthfinder.gov/
- 25. Murphy J. KidsHealth< www. kidshealth. org>. J Consum Health Internet. 2018;22:362-370. [Google Scholar]
- 26. Izacard G, Caron M, Hosseini L, et al. Unsupervised dense information retrieval with contrastive learning. In: Transactions on Machine Learning Research. 2022. https://openreview.net/forum?id=jKN1pXi7b0
- 27. Asai A, Schick T, Lewis P, et al. Task-aware retrieval with instructions. Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics; 2023:3650-3675.
- 28. Su D, Li X, Zhang J, et al. Read before Generate! Faithful Long Form Question Answering with Machine Reading. Findings of the Association for Computational Linguistics: ACL 2022. Association for Computational Linguistics; 2022:744-756.
- 29. Zong C, Wan J, Tang S, et al. EvidenceMap: learning evidence analysis to unleash the power of small language models for biomedical question answering. CoRR 2025. [DOI] [PubMed]
- 30. Meta A. Introducing llama 3.1: our most capable models to date, 2024. Accessed September 2025. https://ai.meta.com/blog/meta-llama-3-1/
- 31. Tiedemann J, Thottingal S. OPUS-MT—Building open translation services for the World. Annual Conference of the European Association for Machine Translation. European Association for Machine Translation; 2020:479-480.
- 32. Junczys-Dowmunt M, Heafield K, Hoang H, et al. Marian: cost-effective high-quality neural machine translation in C++. In: Proceedings of the 2nd Workshop on Neural Machine Translation and Generation. Association for Computational Linguistics; 2018:129-135. http://aclweb.org/anthology/W18-2716
- 33. Tiedemann J. News from OPUS—A collection of multilingual parallel corpora with tools and interfaces. Recent Advances in Natural Language Processing V. John Benjamins Publishing Company; 2009:237-248. [Google Scholar]
- 34. Tiedemann J, Aulamo M, Bakshandaeva D, et al. Democratizing neural machine translation with OPUS-MT. Lang Resour Eval. 2024;58:713-755. [Google Scholar]
- 35. Nakatani S. Language detection library for java, 2010. Accessed January 2025. https://github com/shuyo/language-detection
- 36. Yellamma P, Varun PR, Narayana NCNL, et al. Automatic and multilingual speech recognition and translation by using Google Cloud API. 2024 5th International Conference on Mobile Computing and Sustainable Informatics (ICMCSI). IEEE; 2024:566-571.
- 37. Wang L, Yang N, Huang X, et al. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:221203533. 2022, preprint: not peer reviewed.
- 38. McInnes L, Healy J, Saul N, et al. UMAP: uniform manifold 1120 approximation and projection. JOSS. 2018;3:861. [Google Scholar]
- 39. Campello RJ, Moulavi D, Zimek A, et al. Hierarchical density estimates for data clustering, visualization, and outlier detection. ACM Trans Knowl Discov Data (TKDD). 2015;10:1-51. [Google Scholar]
- 40. Grootendorst M. BERTopic: neural topic modeling with a class-based TF-IDF procedure. arXiv e-prints: arXiv: 2203.05794, preprint: not peer reviewed, 2022.
- 41. Blei DM, Ng AY, Jordan MI. Latent Dirichlet allocation. J Machine Learning Res. 2003;3:993-1022. [Google Scholar]
- 42. Amil S, Da S-M-A-R, Plaisimond J, et al. Interactive conversational agents for perinatal health: a mixed methods systematic review. Healthcare. 2025;13:363. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Hussain S, Ameri Sianaki O, Ababneh N. A survey on conversational agents/chatbots classification and design techniques. Workshops of the International Conference on Advanced Information Networking and Applications. Springer; 2019:946-956.
- 44. Palanica A, Flaschner P, Thommandram A, et al. Physicians’ perceptions of chatbots in health care: cross-sectional web-based survey. J Med Internet Res. 2019;21:e12887. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Elmasri D, Maeder A. A conversational agent for an online mental health intervention. International Conference on Brain and Health Informatics. Springer; 2016:243-251.
- 46. Kowatsch T, Volland D, Shih I, et al. Design and evaluation of a mobile chat app for the open source behavioral health intervention platform MobileCoach. International Conference on Design Science Research in Information System and Technology. Springer; 2017:485-489.
- 47. Murali P, O’Leary T, Shamekhi A, et al. Health counseling by robots: modalities for breastfeeding promotion. 2019 28th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE; 2019:1-6.
- 48. Stroop A, Stroop T, Zawy Alsofy S, et al. Large language models: are artificial intelligence-based chatbots a reliable source of patient information for spinal surgery? Eur Spine J. 2024;33:4135-4143. [DOI] [PubMed] [Google Scholar]
- 49. Fitzpatrick KK, Darcy A, Vierhile M. Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Ment Health. 2017;4:e7785. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. ShojaeiBaghini M, Bahaadinbeigy K. Artificial intelligence-powered chatbot intervention for maternal health: a systematic review study. J Res Health. 2025;15:437-446. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The de-identified versions of data sets generated during and/or analyzed during this study are available from the corresponding author upon reasonable request.





