Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Dec 01, 2023)
Date Submitted: Oct 6, 2022
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Extracting New Meaning from Scientific Abstracts: Combining Linguistic Analysis and Traditional Text Clustering Algorithms in Gaining Novel Insights
ABSTRACT
Background:
Theme extraction by way of topic modeling has increasingly become a central interest in a variety of disciplines seeking to better understand large, unstructured text datasets without the burden of manual text analysis or the need for mechanisms and language experts to make sense of those analyses. Its popularity has resulted in the development of numerous algorithmic approaches and a revolutionary view on the value of studying natural language in many different contexts. Although these machine learning methods have brought a wealth of knowledge, there remains a certain degree of elusiveness in defining the language processing that occurs therein and therefore in the interpretability of its product.
Objective:
This study sought to capitalize on the synergy of linguistic analysis and traditional computational models to extend the capabilities of topic modeling, providing more refined and accessible results that will enhance our understanding of thematic elements and their relationships.
Methods:
To demonstrate this synergy, a corpus of over 22,000 abstracts from the American Society of Anesthesiologists Annual Meeting (ASAAM) between the years 2000-2013 was used. The corpus was analyzed using our novel method, “model of models” (MoM), which encompasses a blend of information retrieval, machine learning, sophisticated data visualizations and human verification. Then, these outputs were compared to both a gold standard of topic classification (i.e., ASAAM’s manual organization of accepted abstracts) and a classical clustering approach.
Results:
Our novel analytical method, MoM, indicated a high level of agreement with the gold standard of ASAAM classification. The MoM’s information-rich yet interpretable network visualization produced an interactive way of exploring the subtle relationships existent in the model, including cluster to cluster and cluster to model relationships, allowing for straightforward verification of validity and coherence as well as convenient interpretation of its results. The clusters produced were logically oriented, reflective of the abstracts’ latent connections with one another, and mostly paralleled those designated by the ASAAM. There were some gaps between the findings of the MoM approach and the traditional text clustering approach, resulting in discrepancies in terms of the best performing model or algorithm.
Conclusions:
This study illustrated the value of re-introducing human input into our machine learning methodologies. Our findings demonstrate an extended functionality of tried-and-true topic modeling approaches, maintaining the computational efficiency offered through traditional techniques while adding a delicate nuance that allows for a deeper exploration of extracted themes, a better understanding of why those themes and their corresponding relationships exist, and the ability to verify the coherence of machine output. The study also reveals the disparities between our proposed method and classic techniques, opening a scholarly dialogue for future refinements of theme extraction and a more effective way to integrate linguistic analysis in an external evaluation process.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.