Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 6, 2026
Date Accepted: Sep 16, 2026
Efficiency vs. Depth: A Comparative Study of AI-generated and Human-synthesized Qualitative Analyses of Ugandan Women’s Experiences of Obstetric Fistula
ABSTRACT
Background:
Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations.
Objective:
Our research sought to (1) describe the process of using artificial intelligence (AI) to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research.
Methods:
Nested within a larger mixed-methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semi-structured interviews exploring their mental, physical, and overall health trajectories. Each interview lasted about an hour and was audio recorded. Following translation and transcription, the data underwent human analysis and analysis using Versa, a University of California San Francisco (UCSF) developed LLM powered by ChatGPT-4o. The AI analysis was conducted using two strategies: inductively (without a codebook) and deductively (using the human-developed codebook). Finally, the outputs were compared to evaluate the code frequency and alignment, thematic depth, analysis quality, and efficiency of each method.
Results:
A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged. When inductively coding, Versa identified 39 codes, whereas human researchers utilized a more expansive set of 54 codes. While Versa’s thematic analyses were generally accurate, human-synthesized themes were more robust. The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.5 hours.
Conclusions:
While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat. Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences present in qualitative data. Further, the misrepresentation of participant excerpts as direct quotes by Versa presents a significant threat to research integrity and ethics. While Versa and similar LLM models can serve as a powerful assistant for efficiency, human-driven analysis is essential to maintain ethical and interpretive rigor. Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed. Clinical Trial: ClinicalTrials.gov NCT05437939; https://clinicaltrials.gov/study/NCT05437939
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.