Previously submitted to: JMIR Medical Informatics (no longer under consideration since Apr 15, 2026)
Date Submitted: Dec 5, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Machine Learning-Based Auxiliary Review of Rational Medication Use in Stress-Related Mucosal Disease: A Clinical Natural Language Text Classifier for Chinese Scenarios
ABSTRACT
Background:
Machine learning has been extensively applied in medical prediction. However, existing machine-driven rational drug use review systems predominantly rely on codifiable information while neglecting unstructured data such as electronic health records (EHRs) documented in natural language, leading to the oversight of potential medication-related risks.
Objective:
To address the aforementioned limitations, this study proposes and validates a text classification-based machine learning framework for the automatic identification of "prophylactic" and "non-prophylactic" medication indications in electronic health records (EHRs), thereby facilitating the assessment of the rationality of prophylactic proton pump inhibitor (PPI) use.
Methods:
In this study, a text annotation approach was adopted, wherein clinical pharmacists performed paragraph-by-paragraph annotation on 3695 medical record segments. Six algorithms—logistic regression, support vector machine (SVM), random forest, decision tree, LightGBM, and CatBoost—were utilized for model training. To address the class imbalance issue in the dataset, SMOTE (Synthetic Minority Oversampling Technique) was employed for oversampling to balance the sample distribution. The model performance was evaluated on an independent test set with an 8:2 training-to-test split ratio. The trained model was then used to validate the test set, predicting whether the text content was indicative of preventive intent or non-preventive intent, thereby determining the rationality of the clinical prophylactic use of proton pump inhibitors (PPIs). Finally, the SHAP (SHapley Additive exPlanations) method was applied to interpret the model's decision-making process.
Results:
The results indicated that all six trained machine learning models exhibited satisfactory performance in the predictive classification of text. Among these models, LightGBM and CatBoost, following SMOTE oversampling, achieved the optimal performance.
Conclusions:
In conclusion, the natural language processing (NLP) text classifier trained based on machine learning can effectively capture medication indications in medical records, providing high-reliability decision support for rational drug use review. Clinical Trial: Not Applicable.This study is a retrospective data analysis based on electronic health records, without involving human subjects or interventional clinical trials, thus trial registration is not applicable.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.