Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 5, 2026
Date Accepted: Aug 18, 2026
The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement from Clinical Decision-making Quality
ABSTRACT
Background:
Large language models (LLMs) have been proposed as decision-support tools in assisted reproductive technology (ART), but it remains unclear whether different post-training alignment strategies translate into clinically acceptable decision support. Outcome-based benchmarks may reward token-level correctness while overlooking the reasoning quality that clinicians rely on.
Objective:
To evaluate whether four mainstream alignment paradigms for medical LLMs: supervised fine-tuning (SFT), direct preference optimization (DPO), group relative policy optimization (GRPO), and in-context learning (ICL), produce comparable algorithmic and clinical alignment (SFT vs GRPO) when used for infertility diagnosis and treatment planning.
Methods:
This retrospective single-center study used 8,201 de-identified electronic health records from West China Second University Hospital, collected between January 2020 and December 2022 (mean age 31.79 years, SD 4.63). All four alignment strategies were built on a shared open-source biomedical backbone. A dual evaluation framework was applied: (1) automatic field-level metrics (accuracy, macro-F1, mean absolute error) on five structured decision fields (infertility type, initial diagnosis, ART strategy, controlled ovarian stimulation (COS) regimen, gonadotropin starting dose); and (2) blinded independent expert review by two reproductive medicine specialists on 100 paired cases across four clinical dimensions (reasoning capability, diagnostic accuracy, treatment feasibility, hallucination). Automatic field-level evaluation included all four alignment strategies, whereas blinded expert review was restricted to the clinically most informative SFT-versus-GRPO contrast.
Results:
GRPO achieved the highest average automatic performance (e.g., Infertility Type accuracy 92.57%; COS regimen accuracy 62.36%; ART strategy Accuracy 76.49%; Gn dose MAE 44.94). However, in blinded expert review, the conservative SFT baseline showed directionally higher expert ratings than GRPO on reasoning capability and treatment feasibility; diagnostic-accuracy differences were not significant. In the three-way best-response comparison including the original physician-charted plan, the SFT baseline was selected as the best response in 51.2% of cases compared with 26.2% for GRPO and 22.6% for the charted plan. This best-response result should be interpreted as a preference within the standardized review format rather than as evidence that model-generated decisions are clinically superior to physician decision-making. Hallucination rates were 15.0% (GRPO) and 18.5% (SFT), indicating that higher automatic performance did not eliminate clinically unsupported content and that neither model is ready for clinical deployment. Subgroup analyses showed GRPO improved F1 in IVF and preimplantation genetic testing (PGT) but decreased F1 in intracytoplasmic sperm injection (ICSI) cases, where male-factor information was largely captured only in unstructured fields.
Conclusions:
Outcome-based metrics alone are insufficient proxies for clinical utility in ART decision support. Algorithmic improvement and clinical alignment suggest a possible decoupling; a phenomenon we term the Alignment Paradox. However, because the clinical review was based on two reproductive medicine specialists with marginal inter-rater agreement, the clinical-alignment findings should be interpreted as exploratory. External multi-center validation is required before clinical deployment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.