Accepted for/Published in: JMIR Formative Research
Date Submitted: Mar 22, 2026
Date Accepted: Sep 8, 2026
Comparative Evaluation of ChatGPT, DeepSeek, and Physician-Generated Acupuncture Treatment Protocols Across Chinese and English Language Settings: A Cross-Sectional Evaluation Study
ABSTRACT
Background:
As an important pillar of integrative medicine, Traditional Chinese Medicine (TCM) has been receiving increasing attention within the field of evidence-based medicine. However, due to its largely experience-based and practice-oriented nature, the process of modernization and standardization has progressed relatively slowly. Meanwhile, rapid advances in artificial intelligence (AI) have created new opportunities for innovation, making the integration of AI technologies with TCM an emerging and promising direction for future development.
Objective:
This study evaluates the application of large language models (LLMs), including ChatGPT and DeepSeek, in generating acupuncture treatment protocols. Specifically, we compare the validity and clinical effectiveness of physician-generated protocols with those produced by LLMs. We also examine differences in outputs across language corpora (Chinese and English) to assess the impact of linguistic context on LLM performance. Ultimately, this research aims to support the standardization of acupoint selection and promote greater consistency, reliability, and evidence-based development in acupuncture practice.
Methods:
Qualified acupuncture treatment cases published in Acupuncture in Medicine were identified and translated from English into Chinese to create a bilingual case dataset. These standardized case prompts were independently input into two LLMs, DeepSeek and ChatGPT, to generate acupuncture treatment protocols. The validity and clinical effectiveness of the AI-generated protocols were evaluated by senior TCM physicians. Treatment protocols produced under different language corpora (Chinese and English) were also compared with physician-generated protocols from the original cases. Evaluations were conducted using a newly developed assessment framework designed for this study, which included the following criteria: local point selection, distal point selection, syndrome-based point selection, meridian-tracing point selection, neuroanatomical point selection, core acupoint selection, and synergistic effects. Statistical analysis was performed using a repeated-measures ANOVA framework with sphericity assessment to examine differences across LLMs, language corpora, and protocol types.
Results:
For local point selection principle, DeepSeek Chinese is higher than ChatGPT Chinese (mean difference = .2, p = .022), ChatGPT English (mean difference = .157, p = .021), and the journal-reported case outcome (mean difference = .314, p = .012). ChatGPT English is higher than ChatGPT Chinese (mean difference = .2, p =.022). For syndrome-based point selection principle, the journal-reported case outcome is lower than DeepSeek Chinese (mean difference = .314, p = .026), DeepSeek English (mean difference = .293, p = .043), and ChatGPT Chinese (mean difference = .414, p = .008). For neuroanatomical point selection principle, DeepSeek English is higher than ChatGPT English (mean difference = .229, p <.001), ChatGPT Chinese (mean difference = .171, p = .002), DeepSeek Chinese (mean difference =.107, p =.038), and the journal-reported case outcome (mean difference = .336, p = .002). For core acupoints coverage principle, the journal-reported case outcome is lower than DeepSeek Chinese (mean difference = .207, p = .028), DeepSeek English (mean difference = .300, p = .005), and ChatGPT English (mean difference = .186, p = .019). The principles of distal point selection, meridian-tracing point selection, and therapeutic synergy were not significantly reflected in the differences among the five groups of the acupuncture prescription (DeepSeek-English, DeepSeek-Chinese, ChatGPT-English, ChatGPT-Chinese, and the journal-reported case outcome). The results indicated clear performance differences across LLMs and language contexts. Overall, DeepSeek demonstrated superior performance in the Chinese-language setting, whereas ChatGPT performed better in the English-language setting. Notably, DeepSeek-generated acupuncture treatment protocols in English achieved the highest scores for neuroanatomical point selection. In contrast, the treatment protocols developed by TCM physicians in the original Acupuncture in Medicine cases received the lowest overall scores across the evaluated criteria.
Conclusions:
These findings suggest that artificial intelligence has made substantial progress in encoding explicit acupuncture knowledge and performing systematic acupoint selection. However, notable challenges remain in delivering individualized treatment and ensuring consistency across linguistic contexts. Further enhancement of bilingual/ multilingual training corpora and the development of syndrome-specific reasoning modules may significantly improve the applicability and clinical utility of AI-assisted systems in TCM practice.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.