Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: Journal of Medical Internet Research

Date Submitted: Nov 4, 2025
Date Accepted: May 5, 2026

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/86453

The final accepted version (not copyedited yet) is in this tab.

“Small” Large Language Models in the Hospital: An Evaluation Study on Real-World Data in a Resource-Constrained Setting

  • He A. Xu; 
  • Romain Pythoud; 
  • Christian W Thorball; 
  • Giorgia Carra; 
  • Bogdan Kulynych; 
  • Jérémie Despraz; 
  • Coralie Galland-Decker; 
  • Errikos Maslias; 
  • Edouard Baudson; 
  • Thomas Brahier; 
  • Vanessa Kraege; 
  • Ana Catarina de Sousa Teixeira; 
  • Carlos Fidalgo; 
  • Florian Berthaudin; 
  • Amagoia Madina; 
  • Solange Zoergiebel; 
  • Francois Bastardot; 
  • Athina Stravodimou; 
  • Marie Méan; 
  • Jacques Fellay; 
  • Philippe Ryvlin; 
  • Jean Louis Raisaro

ABSTRACT

Background:

Large Language Models (LLMs) offer promise for healthcare but face challenges of scale, privacy, and limited evidence in non-English settings. Smaller, locally deployable LLMs remain underexplored.

Objective:

To assess the feasibility of small open-source LLMs (1–24B parameters) in clinical tasks and provide a reproducible hospital-based evaluation framework.

Methods:

Six state-of-the-art small LLMs from the Mistral, Phi-4, Llama-3.1, Meditron-3, Falcon 3 model families were tested in a zero-shot setting on de-identified discharge letters across seven use cases, including information extraction, translation from French to English, summarization, and clinical decision support. Performance was measured with F1 scores, readability indices, embedding similarity, and structured clinician reviews.

Results:

The models achieved high recall in simple retrieval tasks (up to 99.6%) but showed poor performance in protected health information detection, immune-related adverse events detection, summarization, and decision support. Translation quality varied, with general-purpose models outperforming medical-focused models.

Conclusions:

In localized resource-constrained deployments, small LLMs are suitable for basic tasks but insufficient for complex reasoning or clinical decision-making. Our framework supports context-specific evaluation for safe adoption in hospitals.


 Citation

Please cite as:

Xu HA, Pythoud R, Thorball CW, Carra G, Kulynych B, Despraz J, Galland-Decker C, Maslias E, Baudson E, Brahier T, Kraege V, de Sousa Teixeira AC, Fidalgo C, Berthaudin F, Madina A, Zoergiebel S, Bastardot F, Stravodimou A, Méan M, Fellay J, Ryvlin P, Raisaro JL

“Small” Large Language Models in the Hospital: An Evaluation Study on Real-World Data in a Resource-Constrained Setting

Journal of Medical Internet Research. 05/05/2026:86453 (forthcoming/in press)

DOI: 10.2196/86453

URL: https://preprints.jmir.org/preprint/86453

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.