Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Apr 27, 2026
Date Accepted: Jul 15, 2026
Date Submitted to PubMed: Jul 15, 2026

The final, peer-reviewed published version of this preprint can be found here:

Effects of Model Choice, Corpus Context, and Post Hoc Correction on Layer-Level Embedding Degradation in Clinical Document Retrieval: Experimental Study

Mikkelsen Y

Effects of Model Choice, Corpus Context, and Post Hoc Correction on Layer-Level Embedding Degradation in Clinical Document Retrieval: Experimental Study

JMIR Med Inform 2026;14:e99639

DOI: 10.2196/99639

PMID: 42455615

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Layer-Level Analysis of Embedding Degradation in Clinical Document Retrieval: Disentangling Model Tuning, Context, and Recoverability

  • Yngve Mikkelsen

ABSTRACT

Background:

Retrieval-augmented generation (RAG) systems use embedding models for clinical documents. A companion study found that domain-specific encoders (BioBERT, ClinicalBERT) underperform general-purpose embeddings and produce near-degenerate embedding geometry, but did not pinpoint the source of the failure or whether it is recoverable without retraining.

Objective:

We aimed to (a) identify transformer layers in which embedding degradation emerges; (b) determine whether degradation reflects model tuning or input context; (c) characterize how their relative contributions change with network depth; and (d) evaluate whether post-hoc interventions can restore performance without retraining.

Methods:

Embeddings were extracted from every hidden layer of 11 transformer configurations across 3 clinical corpora (n=100 each: MTSamples, PMC-Patients, Mistral-7B-Instruct-v0.2-generated synthetic notes) and 2 query formats. GPT-4o generated metadata queries; synthetic alignment for mixed-effects analysis used BM25 deduplication. Embedding geometry (participation ratio, average pairwise cosine) and retrieval performance (MRR@10, Recall@10) were measured. A Type II ANOVA estimated per-layer factor contributions; a per-query linear mixed-effects model on final-layer BERT-scale data served as a sensitivity check for non-independence. Four no-retraining interventions were assessed: layer selection, layer combination, zero-phase component analysis (ZCA) whitening, and mean centering. Generalization was tested on validation corpora (500 PMC-Patients, 400 MTSamples).

Results:

All 11 configurations exhibited a U-shaped retrieval curve with mid-layer collapse (mean normalized MRR = 0.04 ± 0.05 at 35–55% relative depth versus 0.86 at the final layer). General-purpose models increased effective dimensionality by 2.3–4.0× from trough to final layer, compared with 1.0–1.4× for domain encoders. In Type II ANOVAs at each BERT-scale layer, model η² declined from 0.85 at layer 0 to 0.70 at the final layer, while corpus η² rose from 0.01 to 0.24, indicating that context grows in influence with depth. Transductive ZCA whitening yielded the largest post hoc gains (mean MRR@10 +0.19 BioBERT, +0.27 E5-Mistral-7B, +0.27 Phi-3-mini; upper bound). The deployment-realistic corpus-only variant was positive across all 16 natural-text conditions but harmful across all 8 synthetic-corpus conditions. Cross-source model rankings between case reports and medical transcriptions were nearly identical (Spearman ρ = 0.976, p < .001).

Conclusions:

Mid-layer embedding collapse was observed across all 11 examined transformer configurations; high final-layer performance was concentrated among retrieval-oriented, contrastively trained models, consistent with retrieval-oriented training rather than domain pretraining. A per-query linear mixed-effects sensitivity analysis of final-layer BERT-scale data confirmed that model and corpus effects dominated query format (model: χ²=2996.4, df=7, p<.001; corpus: χ²=88.2, df=2, p<.001; query format: χ²=0.05, df=1, p=.83), with ICC=0.380 and marginal pseudo-R²=0.389 closely matching the companion study (pseudo-R²=0.373). Model rankings remained stable across two independent validation corpora. We propose a practical pre-deployment evaluation protocol covering layer selection, embedding-geometry diagnostics, and conditional post-hoc correction, pending validation with institutional electronic health record data.


 Citation

Please cite as:

Mikkelsen Y

Effects of Model Choice, Corpus Context, and Post Hoc Correction on Layer-Level Embedding Degradation in Clinical Document Retrieval: Experimental Study

JMIR Med Inform 2026;14:e99639

DOI: 10.2196/99639

PMID: 42455615

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.