Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Informatics

Date Submitted: Aug 2, 2026
Open Peer Review Period: Aug 5, 2026 - Sep 30, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

NTCIR-18 RadNLP 2024 Overview: Dataset and Solutions for Automated Lung Cancer Staging

  • Yuta Nakamura; 
  • Eiji Aramaki; 
  • Shuntaro Yada; 
  • Koji Fujimoto; 
  • Shouhei Hanaoka; 
  • Jonas Kluckert; 
  • Michael Krauthammer; 
  • Jun Kanzawa; 
  • Akira Katayama; 
  • Tomohiro Kikuchi; 
  • Ryo Kurokawa; 
  • Wataru Gonoi; 
  • Yuki Tashiro

ABSTRACT

Background:

Automated understanding of radiology reports has become increasingly important for clinical decision support, reg- istry construction, and the development of large language models (LLMs). However, most benchmark datasets are limited to English, hindering progress in multilingual clinical natural language processing (NLP).

Objective:

We present the NTCIR-18 RadNLP 2024 benchmark, a multilingual shared task for automated lung cancer TNM staging from radiology reports. Building upon three consecutive NTCIR editions, the benchmark provides Japanese and English radiology reports together with standardized evaluation protocols and community baselines.

Methods:

The NTCIR-18 RadNLP 2024 task attracted 19 participating teams from academia and industry. Participants developed systems using approaches ranging from conventional machine learning to state-of-the-art LLMs. We summarize the dataset, task design, participating methods, and benchmark results.

Results:

LLM-based methods substantially outperformed conventional approaches, although performance varied across TNM categories and languages. The benchmark also highlights persistent challenges in clinically faithful information extraction from radiology reports.

Conclusions:

Rather than serving solely as a shared-task report, the NTCIR-18 RadNLP 2024 benchmark establishes a reusable evaluation resource for multilingual clinical NLP. We anticipate that it will facilitate future research on structured information extraction, report generation, and clinically grounded evaluation of radiology LLMs.


 Citation

Please cite as:

Nakamura Y, Aramaki E, Yada S, Fujimoto K, Hanaoka S, Kluckert J, Krauthammer M, Kanzawa J, Katayama A, Kikuchi T, Kurokawa R, Gonoi W, Tashiro Y

NTCIR-18 RadNLP 2024 Overview: Dataset and Solutions for Automated Lung Cancer Staging

JMIR Preprints. 02/08/2026:107996

DOI: 10.2196/preprints.107996

URL: https://preprints.jmir.org/preprint/107996

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.