Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Public Health and Surveillance

Date Submitted: Apr 27, 2026
Date Accepted: Aug 8, 2026

The final, peer-reviewed published version of this preprint can be found here:

Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

Luo Z, Epstein RA, Sambamoorthi N, Muhammad LN, Jordan N

Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

JMIR Public Health Surveill 2026;12:e99643

DOI: 10.2196/99643

PMID: 42748426

State child maltreatment policies and maltreatment outcomes: A multi-state study using natural language processing

  • Zhidi Luo; 
  • Richard A Epstein; 
  • Nethra Sambamoorthi; 
  • Lutfiyya N Muhammad; 
  • Neil Jordan

ABSTRACT

Background:

Child maltreatment is a major public health issue in the United States, with substantial variation in how states define, report, and respond to abuse and neglect. While prior research has examined individual policy components, less is known about how multiple policies co-occur to form broader policy environments and how these multidimensional configurations relate to maltreatment outcomes.

Objective:

This study aimed to use natural language processing (NLP) to systematically characterize state child maltreatment policy environments and examine their associations with maltreatment incidence, recurrence, and fatalities across U.S. jurisdictions.

Methods:

A cross-sectional study was conducted using 2021 data from all 50 U.S. states, the District of Columbia, and Puerto Rico (N = 52). State maltreatment policies were derived from the State Child Abuse & Neglect Policies Database, comprising 411 policy items spanning definitions, reporting, screening, investigation, response, and system context. Six NLP models (BART, BERT, RoBERTa, DeBERTa, Copilot, and LLaMA 3.1) were applied using a zero-shot classification framework to quantify policy characteristics. Models were evaluated using intrinsic (category consistency and semantic alignment) and extrinsic (factor analysis and clustering performance) metrics. Exploratory factor analysis (EFA) was used to identify latent policy domains, and k-means clustering was applied to group jurisdictions with similar policy profiles. Maltreatment outcomes, including incidence, recurrence, and fatalities, were obtained from the National Child Abuse and Neglect Data System. Outcome differences across clusters were assessed using analysis of variance with post-hoc pairwise comparisons.

Results:

The study population included 72,838,819 children and 3,774,528 maltreatment reports in 2021, of which 751,283 were indicated cases. Across NLP model evaluations, DeBERTa demonstrated the best overall performance. EFA identified three primary policy domains: maltreatment definition, mandated reporting, and alternative response. Four distinct policy clusters were identified. Jurisdictions with weaker reporting requirements and fewer penalties (Cluster 2) had the highest incidence of maltreatment (mean 17.51 vs. 9.65–11.74 per 1,000 children; p = 0.069) and recurrence (1.74 vs. 0.52–0.68 per 1,000 children; p < 0.001). In contrast, jurisdictions with stronger reporting requirements, broader definitions, and greater use of alternative response systems (Cluster 1) had the lowest incidence and recurrence. No statistically significant differences in maltreatment fatalities were observed across clusters (p = 0.889). However, in a sub-analysis of 18 fatality-related policy items, clusters differed significantly in fatality rates (p = 0.029), with jurisdictions characterized by less clearly defined fatality policies exhibiting higher fatality rates (47.91 vs. 19.03–25.63 per 1,000,000 children; p = 0.029).

Conclusions:

State child maltreatment policies form distinct, multidimensional configurations that are differentially associated with maltreatment outcomes. The findings highlight the importance of evaluating integrated policy frameworks rather than isolated components and demonstrate the utility of NLP for large-scale, data-driven policy analysis.


 Citation

Please cite as:

Luo Z, Epstein RA, Sambamoorthi N, Muhammad LN, Jordan N

Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

JMIR Public Health Surveill 2026;12:e99643

DOI: 10.2196/99643

PMID: 42748426

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.