Medicine

ChatGPT-4o Shows Promise but Makes Critical Errors Translating Medical Conversations in Nepal

How the science connects

Natural language p…Machine translation

AI Insight

This field study evaluated ChatGPT-4o's accuracy in real-time English-Nepali voice translation during conversations with 30 Nepali-speaking adults in rural Nepal. Overall, 58% of 485 translations received the highest accuracy rating, but performance was significantly worse for Nepali-to-English translation (91% of low-accuracy translations occurred in this direction). Common errors included distortion of intended meaning, unnatural phrasing, omissions, and additions of content not present in the original speech.


This research addresses the practical utility of AI translation tools in under-resourced language settings where trained interpreters are scarce, relevant for community health research and clinical communication. The findings highlight that while ChatGPT-4o shows promise for basic conversational exchange, errors substantial enough to alter meaning require human verification in research or clinical contexts where accuracy is critical.


Understand the Science

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.

Source: Accuracy and error patterns of ChatGPT-4o for real-time English-Nepali voice translation: A cross-sectional field evaluation in rural Nepal