Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 11, 2026
Open Peer Review Period: Aug 12, 2026 - Oct 7, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Cross-Site Stability of Embedding Model Rankings for Known-Item Retrieval From Clinical Notes: A Two-Corpus Comparative Evaluation Study

  • Yngve Mikkelsen

ABSTRACT

Background:

Retrieval-augmented generation over clinical text depends on the embedding model chosen, and this choice interacts with documentation genre. Cross-site stability of embedding-model rankings has not been evaluated on real EHR narratives with genre held constant; public benchmarks are predominantly single-source, and multi-institution evaluations address within-task matching rather than ranking stability.

Objective:

To test whether embedding model rankings transfer across two institutional EHR corpora with genre held constant, and to examine the contributions of model, site/corpus, and genre to retrieval effectiveness and how they depend on scoring choices.

Methods:

We compared 13 embedding models on a known-item retrieval task using a 2×2 design that crossed site/corpus (Beth Israel Deaconess Medical Center via MIMIC-IV-Note; University of California San Francisco via ER-Reason) with documentation genre (discharge summary; diagnostic imaging report). To prevent verbatim query-target overlap, we removed each query's sentences from its target before indexing; deduplicated notes after normalization; used model-appropriate pooling configurations; scored long documents as overlapping chunks using maximum chunk similarity; and matched sample sizes across cells (N=1235 per cell). Retrieval effectiveness was measured using mean reciprocal rank at 10 (MRR@10); cross-site rank transfer using Kendall τ and selection regret; and variance using factorial decomposition (η²). Transfer and regret estimates and the contrastive-panel decomposition carry 95% patient-clustered bootstrap intervals.

Results:

Model rankings transferred across sites within both genres: among the 8 contrastive models, cross-site Kendall τ was 0.86 for discharge and 0.71 for imaging (95% patient-clustered CIs 0.64–0.93 and 0.43–0.93). The model×site interaction was small under every scoring choice (η²=0.010–0.028); the site/corpus main effect was negligible under primary best-chunk scoring (η²=0.005) but scoring-dependent (0.160 under mean-pooling). Observed cross-site selection regret was 0 for imaging and up to 0.041 for discharge (~12% of the destination's best MRR; bootstrap intervals 0–0.083). Model and genre were major sources of variation (contrastive-panel η²=0.33 and 0.54); their ordering was scoring-sensitive, but genre remained major after chunk-count and anisotropy checks. Adding 5 nonretrieval MLM comparators moved the largest term to model (η²=0.87), showing that the dominant-factor question is panel-sensitive, while transfer was not.

Conclusions:

For known-item retrieval from clinical notes, embedding-model rankings showed moderate-to-strong cross-site concordance (Kendall τ 0.71–0.86 among contrastive models). The model×site interaction remained small across scoring choices, whereas the site/corpus main effect was scoring-dependent. Cross-site model selection incurred zero to modest performance loss. Model choice and documentation genre were major determinants of absolute effectiveness. These findings suggest that, when documentation genre and retrieval configuration are matched, model-selection results from one institutional EHR corpus may provide a useful starting point for another, although local confirmation remains warranted.


 Citation

Please cite as:

Mikkelsen Y

Cross-Site Stability of Embedding Model Rankings for Known-Item Retrieval From Clinical Notes: A Two-Corpus Comparative Evaluation Study

JMIR Preprints. 11/08/2026:109305

DOI: 10.2196/preprints.109305

URL: https://preprints.jmir.org/preprint/109305

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.