Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Jan 5, 2026
Date Accepted: Aug 14, 2026

The final, peer-reviewed published version of this preprint can be found here:

Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act–Aligned Framework

sblendorio E, Dentamaro V, De Maria M, Tempesta S, Barile E, Napolitano D, Tallini M, Nigrelli D, Cicolini G, Piredda M

Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act–Aligned Framework

JMIR Med Inform 2026;14:e90854

DOI: 10.2196/90854

Cloud-Based and Locally Deployed Language Models in Nursing and Healthcare: An AI Act-Aligned Framework

  • Elena sblendorio; 
  • Vincenzo Dentamaro; 
  • Maddalena De Maria; 
  • Salvatore Tempesta; 
  • Elena Barile; 
  • Daniele Napolitano; 
  • Martina Tallini; 
  • Daniela Nigrelli; 
  • Giancarlo Cicolini; 
  • Michela Piredda

ABSTRACT

Background:

The integration of Large Language Models into high-risk systems like healthcare is accelerating. Rigorous evaluations aligned with emerging Legislation are imperative prior to their incorporation into university educational platforms and clinical practice settings.

Objective:

The objective of our study was to implement the first Regulation (EU) 2024/1689-aligned methodological framework for a systematic, comprehensive and dynamically adaptable Language Model evaluation, supporting decision-making in specialized healthcare management.

Methods:

We analyzed 15 Large Language Models (LLMs) and 2 Small Language Models (SLMs). A seven-domain, EU AI Act-aligned methodological framework was employed. Feasibility was tested with a dataset of 32 multiparametric engineered clinical prompts to elicit evaluation in 27 items, with Delphi expert responses as ground truth (all data transparently published). Double-blind interdisciplinary evaluation on a 7-point Likert scale achieved high inter-rater reliability per model (Krippendorff's α = 0.759 on average). A comprehensive analysis identified specific strengths and vulnerabilities. Safety was analyzed as alignment with both Evidence-Based-Nursing and novel structured assessments including ethical resilience testing (via progressive “jailbreaking “). Further novel structured assessments included reference classification, automated consistency and NANDA-I terminology.

Results:

A stringent “Safety-Gatekeeper” domain immediately classified 11 of 17 language models as unsuitable due to critical failures in evidence-based alignment or ethical resilience. GPT-o1, GPT-4o, Gemini 2.0 Pro.Exp, and three Anthropic models surpassed minimum thresholds, permitting evaluation progression. Only Anthropic Sonnet variants achieved uniform “recommended” categorization. For instance, Sonnet 3.7 Thinking produced 75.9% of accurate, focused references, and achieved high average score both in clinical safety and data security (6.73 ± 0.22 and 6.83 ± 0.41, respectively). DeepSeek-R1, Perplexity Sonar, Mistral Large and Qwen2.5-14B failed to resist explicit harmful prompts; Claude 3 Opus resisted even jailbreak attempts, and demonstrated null sycophancy. Notably, Qwen2.5-14B, operating locally, outperformed 4 of the 15 LLMs in multi-step problems in nurse staffing optimization. NANDA-I diagnostic translation capability improved significantly with taxonomy-embedded contexts, with Gemini demonstrating exceptional performance (F1 score 0.59, MAPD = 4.0).

Conclusions:

Regulatory-aligned LLM integration in pilot university hospitals can enhance healthcare education and decision-making across standardized taxonomy, evidence-based personalized clinical care algorithms, computational tasks, and non-discrimination policies, under structured interdisciplinary expert oversight. The methodology demonstrates adaptability across various clinical settings. Future advancements should prioritize multimodal capabilities and locally functioning models, addressing resource disparities in line with Sustainable Development Goal 10, alongside operational resilience, and enhanced data protection. Clinical Trial: Not applicable. This study is a methodological evaluation and does not report health-related outcomes in human participants.


 Citation

Please cite as:

sblendorio E, Dentamaro V, De Maria M, Tempesta S, Barile E, Napolitano D, Tallini M, Nigrelli D, Cicolini G, Piredda M

Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act–Aligned Framework

JMIR Med Inform 2026;14:e90854

DOI: 10.2196/90854

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.