Currently submitted to: JMIR Medical Informatics
Date Submitted: Aug 16, 2026
Open Peer Review Period: Aug 26, 2026 - Oct 21, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
GPT-Assisted Harmonization and Reuse of Cosmetovigilance and Cosmetic-Series Patch-Test Data: Two-Source Retrospective Clinical Informatics Study
ABSTRACT
Background:
Routinely collected dermatology data are valuable for clinical surveillance and secondary research but are often difficult to reuse because of heterogeneous data structures, inconsistent terminology, mixed analytic units, and unclear denominators. Cosmetovigilance reports and cosmetic-series patch-test records provide complementary information on clinical adverse reactions and ingredient-level sensitization; however, these data sources are not readily interoperable. Large language models may facilitate semantic harmonization and structured information extraction, but their use in clinical data processing requires transparent task boundaries and expert oversight.
Objective:
This study aimed to develop and evaluate a human-supervised, GPT-assisted clinical informatics workflow for harmonizing and reusing cosmetovigilance and cosmetic-series patch-test data, with a focus on diagnosis grouping, ingredient-name mapping, reaction-grade standardization, denominator identification, and generation of reusable clinical signals.
Methods:
We conducted a two-source retrospective clinical informatics study using 590 cosmetic-related cutaneous adverse reaction reports submitted between 2017 and 2026 from the West China Hospital sentinel site to the national cosmetovigilance system and records from 485 patients who underwent cosmetic-series patch testing between 2020 and 2024. The two datasets were analyzed as complementary rather than individually linked data sources. A human-curated reference standard was established for diagnosis groups, ingredient mappings, reaction grades, analytic units, and denominators. GPT was applied as an assisted semantic harmonization and information-extraction layer for diagnosis mapping, ingredient-name mapping, reaction-grade standardization, denominator identification, and auditable clinical signal summarization. All GPT-generated outputs were reviewed and validated by domain experts before statistical analysis.
Results:
The workflow resolved the cosmetovigilance dataset into multiple analytic layers, including 590 case-level reports, 987 implicated cosmetic-category records, 988 product-source records, 630 diagnosis records, 1029 symptom records, 647 affected-site records, 1185 lesion-morphology records, and 985 original-product patch-test records. The patch-test dataset comprised 485 patients, 40 cosmetic-series test substances, and 1332 recorded reactions. In the cosmetovigilance data, cosmetic contact dermatitis accounted for 501 of 630 diagnosis records (79.52%), and original-product patch testing was not performed or had unknown status in 675 of 985 records (68.53%). Among 310 completed and interpretable original-product patch tests, 97 were positive (31.29%). In the cosmetic-series cohort, 448 of 485 patients (92.37%) had at least one positive reaction, and methylisothiazolinone was the most frequently positive test substance, with 90 positive reactions among 485 patients (18.56%).
Conclusions:
A GPT-assisted, human-in-the-loop workflow can support the transformation of heterogeneous cosmetovigilance and patch-test data into structured, denominator-resolved, interpretable, and reusable clinical informatics resources. GPT appears most appropriate as an assisted semantic harmonization and information-extraction tool rather than an autonomous diagnostic or causal inference system. Expert validation, de-identification, auditability, and transparent denominator management remain essential for responsible reuse of real-world clinical data. Further multicenter and blinded evaluations are needed to compare GPT-assisted workflows with manual and rule-based approaches in terms of accuracy, efficiency, consistency, and implementation value. Clinical Trial: Not applicable; this was a retrospective clinical informatics study and did not involve a clinical trial.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.