Accepted for/Published in: JMIR Formative Research
Date Submitted: Mar 7, 2026
Date Accepted: Jun 1, 2026
Emotion Classification in Japanese Cancer Survivor Interview Narratives Using Sentiment Polarity and Plutchik’s Emotion Frameworks: Model Development and Evaluation Study
ABSTRACT
Background:
Cancer survivors often experience complex and sometimes coexisting emotions across diagnosis, treatment, and posttreatment life. Capturing these emotions from patient narratives may support clinicians in assessing psychological states and delivering patient-centered psychosocial care, but evidence remains limited on building multi-emotion classifiers from interview narratives. These narratives reflect both personal and social dimensions of survivorship, aligning with the social model of health.
Objective:
This study aimed to develop and evaluate emotion classification models using Japanese cancer survivor interview narratives as a formative step toward narrative-informed support.
Methods:
We analyzed verbatim transcripts from 15 cancer survivor interviews (typically ~90 minutes each) published by the NPO GanNote. From narratives covering diagnosis/disclosure through posttreatment daily life, we preprocessed the transcripts by removing non-informative conversational elements and segmenting text at punctuation into sentence-level chunks, which were then grouped into blocks of five consecutive sentences to fit the 512-token input limit. Two annotators labeled 1,998 text chunks with (1) 3-class sentiment polarity (positive/neutral/negative) and (2) overlapping labels for Plutchik’s eight basic emotions (trust, anticipation, joy, surprise, sadness, fear, disgust, anger). We fine-tuned Japanese BERT and LUKE to build a multi-class polarity classifier and a multilabel eight-emotion classifier, and evaluated performance with F1-scores (including Micro-F1 for polarity and exact match ratio for multilabel). For comparison, we trained the same architectures on WRIME (social media posts with reader-assigned emotions) and tested transfer performance on the interview data.
Results:
The interviews had a mean of 136.9 words per narrative. Label distributions were imbalanced, with the most-to-least frequency ratio of 3.47 (polarity) and 8.10 (eight emotions); neutral and trust were the most frequent labels, whereas anger was the least frequent. Co-occurrence analysis showed sadness and trust most frequently co-occurred, while anger and anticipation co-occurred least. The LUKE model fine-tuned on interview narratives achieved the best performance for polarity classification (Micro-F1=0.73; Macro-F1=0.67; F1: neutral=0.76, positive=0.62, negative=0.63). For eight emotions, the same model achieved Macro-F1=0.42 and exact match ratio=0.51, with higher F1 for trust (0.62), anticipation (0.54), disgust (0.54), sadness (0.52), and fear (0.49) but lower F1 for anger (0.09), joy (0.33), and surprise (0.21). Models trained on WRIME showed lower performance when evaluated on interview narratives (polarity Micro-F1=0.58–0.60; eight-emotion exact match ratio=0.18–0.21).
Conclusions:
NLP-based emotion classification from cancer survivor interview narratives is feasible and domain-specific models outperformed general-purpose classifiers, suggesting the importance of training data that reflects the target narrative context. These findings illustrate that survivorship narratives may contain coexisting and contrasting emotions not fully captured by polarity alone. Addressing class imbalance and improving detection of low-frequency emotions (eg, anger) will be important for future work. Before clinical translation, user-centered evaluation to assess interpretability, safety, and workflow integration will also be needed.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.