Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Dec 30, 2025
Date Accepted: Jun 22, 2026

The final, peer-reviewed published version of this preprint can be found here:

Improving Clinical Validity in Synthetic Electronic Health Record Generation Using Best-of-N Sampling: Comparative Evaluation Study

Masud MA, Hasan M

Improving Clinical Validity in Synthetic Electronic Health Record Generation Using Best-of-N Sampling: Comparative Evaluation Study

JMIR AI 2026;5:e90590

DOI: 10.2196/90590

PMID: 42531495

PMCID: 13422743

Clinical Validity in Synthetic EHR Generation: Best-of-N Sampling in Tabular GANs

  • Md Akmol Masud; 
  • Mahmud Hasan

ABSTRACT

Background:

Synthetic electronic health record generation is limited not only by statistical fidelity but also by clinical validity. Records that appear statistically plausible may still violate hard structural, physiological, or relational constraints.

Objective:

This study evaluated Best-of-N constraint-minimizing selection as an inference-time strategy for improving clinical validity and characterized the conditions under which such selection succeeds or fails as a function of the generator's valid-support mass (pvalid).

Methods:

We evaluated WGAN-GP and CTGAN on three public clinical tabular datasets: stroke prediction (n=5,110), diabetes health indicators (n=100,000), and cardiovascular disease (n=68,599). For each generator, we compared random sampling, naive clipping, and Best-of-N selection (N ∈ {8, 16, 128}). Validity was assessed via dataset-specific rule violations; fidelity via Kolmogorov-Smirnov (KS) statistics and correlation preservation; utility via train-on-synthetic-test-on-real (TSTR) area under the ROC curve (AUC); and privacy via membership inference attack (MIA) AUC and distance to closest record (DCR).

Results:

Best-of-N effectiveness was governed by the base generator's valid-support mass. In stroke, WGAN-GP had pvalid = 0, and all selection strategies—including strict rejection across 5,000 draws—produced 500/500 (100%) invalid samples. In cardiovascular data, WGAN-GP had pvalid = 0.12; Best-of-16 reduced violations from 441/500 (88.2%) to 65/500 (13.0%), and Best-of-128 eliminated them entirely (0/500), matching theoretical predictions exactly. In diabetes, WGAN-GP random sampling was already fully valid (0/500 violations), and Best-of-N primarily shifted utility. Across all three datasets, CTGAN + Best-of-16 achieved 0/500 violations with strong fidelity (mean KS ≤ 0.12) and utility (TSTR AUC up to 0.81). MIA AUC remained near random guessing (approximately 0.50) throughout, though DCR-based proximity concentration increased markedly under CTGAN selection (reaching 22.6% on stroke and 100% on cardiovascular data).

Conclusions:

Best-of-N functions as a bounded-budget feasibility filter rather than a universal repair mechanism. Its effectiveness depends critically on the generator's valid support mass. CTGAN + Best-of-16 offered the strongest overall trade-off across clinical validity, fidelity, utility, and privacy in our experiments.


 Citation

Please cite as:

Masud MA, Hasan M

Improving Clinical Validity in Synthetic Electronic Health Record Generation Using Best-of-N Sampling: Comparative Evaluation Study

JMIR AI 2026;5:e90590

DOI: 10.2196/90590

PMID: 42531495

PMCID: 13422743

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.