Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Apr 27, 2026
Date Accepted: Jul 15, 2026
Date Submitted to PubMed: Jul 15, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Layer-Level Analysis of Embedding Degradation in Clinical Document Retrieval: Disentangling Model Tuning, Context, and Recoverability
ABSTRACT
Background:
Retrieval-augmented generation (RAG) systems use embedding models for clinical documents. A companion study found that domain-specific encoders (BioBERT, ClinicalBERT) underperform general-purpose embeddings and produce near-degenerate embedding geometry, but did not pinpoint the source of the failure or whether it is recoverable without retraining.
Objective:
We aimed to (a) identify transformer layers in which embedding degradation emerges; (b) determine whether degradation reflects model tuning or input context; (c) characterize how their relative contributions change with network depth; and (d) evaluate whether post-hoc interventions can restore performance without retraining.
Methods:
Embeddings were extracted from every hidden layer of 11 transformer configurations across 3 clinical corpora (n=100 each: MTSamples, PMC-Patients, Mistral-7B-Instruct-v0.2-generated synthetic notes) and 2 query formats. GPT-4o generated metadata queries; synthetic alignment for mixed-effects analysis used BM25 deduplication. Embedding geometry (participation ratio, average pairwise cosine) and retrieval performance (MRR@10, Recall@10) were measured. A Type II ANOVA estimated per-layer factor contributions; a per-query linear mixed-effects model on final-layer BERT-scale data served as a sensitivity check for non-independence. Four no-retraining interventions were assessed: layer selection, layer combination, zero-phase component analysis (ZCA) whitening, and mean centering. Generalization was tested on validation corpora (500 PMC-Patients, 400 MTSamples).
Results:
All 11 configurations exhibited a U-shaped retrieval curve with mid-layer collapse (mean normalized MRR = 0.04 ± 0.05 at 35–55% relative depth versus 0.86 at the final layer). General-purpose models increased effective dimensionality by 2.3–4.0× from trough to final layer, compared with 1.0–1.4× for domain encoders. In Type II ANOVAs at each BERT-scale layer, model η² declined from 0.85 at layer 0 to 0.70 at the final layer, while corpus η² rose from 0.01 to 0.24, indicating that context grows in influence with depth. Transductive ZCA whitening yielded the largest post hoc gains (mean MRR@10 +0.19 BioBERT, +0.27 E5-Mistral-7B, +0.27 Phi-3-mini; upper bound). The deployment-realistic corpus-only variant was positive across all 16 natural-text conditions but harmful across all 8 synthetic-corpus conditions. Cross-source model rankings between case reports and medical transcriptions were nearly identical (Spearman ρ = 0.976, p < .001).
Conclusions:
Mid-layer embedding collapse was observed across all 11 examined transformer configurations; high final-layer performance was concentrated among retrieval-oriented, contrastively trained models, consistent with retrieval-oriented training rather than domain pretraining. A per-query linear mixed-effects sensitivity analysis of final-layer BERT-scale data confirmed that model and corpus effects dominated query format (model: χ²=2996.4, df=7, p<.001; corpus: χ²=88.2, df=2, p<.001; query format: χ²=0.05, df=1, p=.83), with ICC=0.380 and marginal pseudo-R²=0.389 closely matching the companion study (pseudo-R²=0.373). Model rankings remained stable across two independent validation corpora. We propose a practical pre-deployment evaluation protocol covering layer selection, embedding-geometry diagnostics, and conditional post-hoc correction, pending validation with institutional electronic health record data.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.