Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Research Protocols

Date Submitted: Apr 29, 2026
Date Accepted: Aug 31, 2026

The final, peer-reviewed published version of this preprint can be found here:

Large Language Models for Clinical Data Extraction in Workplace Injury Rehabilitation: Protocol for a Retrospective Pilot Study of Accuracy, Fairness, and Methodological Considerations in Workers’ Compensation Medical Chart Review

Shah AR, Gorczyca Abel B, Komeili M, Gohar B, Nowrouzi-Kia B

Large Language Models for Clinical Data Extraction in Workplace Injury Rehabilitation: Protocol for a Retrospective Pilot Study of Accuracy, Fairness, and Methodological Considerations in Workers’ Compensation Medical Chart Review

JMIR Res Protoc 2026;15:e99807

DOI: 10.2196/99807

Large Language Models for Clinical Data Extraction in Workplace Injury Rehabilitation: Protocol for a Pilot Study of Accuracy, Fairness, and Methodological Considerations in Workers' Compensation Medical Chart Review

  • Armaan Rehman Shah; 
  • Barbara Gorczyca Abel; 
  • Majid Komeili; 
  • Basem Gohar; 
  • Behdin Nowrouzi-Kia

ABSTRACT

Background:

Return-to-work (RTW) outcomes following workplace injury depend on accurate clinical assessment, timely rehabilitation, and equitable access to care. Electronic medical charts within the Workplace Safety and Insurance Board (WSIB) specialty programs contain rich, but unstructured, clinical and sociodemographic data essential for injury classification, treatment planning, and compensation decisions. Traditional manual chart review is time-consuming, inconsistent, and susceptible to reviewer bias. Large language models (LLMs) have demonstrated strong capabilities in extracting complex information from unstructured clinical text, yet their application to occupational health and workplace injury rehabilitation remains unexplored.

Objective:

This study aims to evaluate the accuracy, fairness, and methodological implications of using LLMs to extract clinical and rehabilitative information from WSIB medical charts at Trillium Health Partners (THP). Specifically, we will assess Artificial Intelligence (AI) models' performance in identifying injury classifications and sociodemographic variables, examine algorithmic bias across demographic groups, and explore the ethical implications of AI-assisted chart review for clinical decision-making and workplace compensation.

Methods:

We will conduct a retrospective chart review of approximately 50 medical records from the WSIB Back and Neck Specialty Program at THP, spanning January 2018 to December 2024. General-purpose models (Qwen3-VL-8B and InternVL3.5-8B) and domain-specific biomedical models (MedGemma-27B and LLaMA-3-Meditron-8B) will be evaluated in their off-the-shelf form using zero-shot and few-shot prompting. In addition, the biomedical models will be fine-tuned on the annotated dataset to assess improvements in structured data extraction performance. To establish ground truth, independent human annotation of a chart subset will be performed, with inter-annotator agreement quantified using Cohen’s kappa. Proposed performance metrics will include sensitivity, specificity, precision, and F1-score for categorical variables; macro-averaged F1-score for multi-class tasks; mean absolute error (MAE) and root mean squared error (RMSE) for continuous variables; and token-level and span-level precision, recall, and F1-score for text extraction tasks. Algorithmic fairness will be examined using Demographic Parity and Equalized Odds across demographic subgroups. The proposed architecture will employ a secure hybrid design in which all personal health information (PHI) will remain within THP on-premise infrastructure, with cloud compute restricted to transient processing.

Results:

This study was funded by the Data Sciences Institute at the University of Toronto in April 2025. Ethics approval will be sought from the Trillium Health Partners and the University of Toronto’s Research Ethics Boards. Future studies seeking ethics approval may build on this protocol to develop data extraction infrastructure, conduct pilot data collection, and share findings on LLM-based chart review in occupational health settings.

Conclusions:

This protocol describes the first systematic evaluation of LLM-based data extraction applied to workplace injury medical charts. Findings will inform best practices for the responsible, reproducible, and equitable integration of AI tools in occupational health clinical workflows and workers' compensation decision-making.


 Citation

Please cite as:

Shah AR, Gorczyca Abel B, Komeili M, Gohar B, Nowrouzi-Kia B

Large Language Models for Clinical Data Extraction in Workplace Injury Rehabilitation: Protocol for a Retrospective Pilot Study of Accuracy, Fairness, and Methodological Considerations in Workers’ Compensation Medical Chart Review

JMIR Res Protoc 2026;15:e99807

DOI: 10.2196/99807

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.