Currently submitted to: JMIR Bioinformatics and Biotechnology
Date Submitted: Jul 17, 2026
Open Peer Review Period: Aug 4, 2026 - Sep 29, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
AI-Guided Synthetic Protein Sequence Generation for Functional Protein Design Using Gated Recurrent Units
ABSTRACT
Background:
Designing functional proteins from first principles confronts a combinatorial barrier of staggering scale: a polypeptide of 100 residues spans 20^100 possible sequences, a search space entirely inaccessible by experimental means alone. This paper presents an end-to-end artificial intelligence pipeline that learns the statistical grammar of proteins directly from the UniProtKB Human Proteome database and generates novel, biologically plausible candidates without structural templates or homologous scaffolds. A character-level Gated Recurrent Unit (GRU) model (64-dimensional embedding, 128 hidden units, 78,356 parameters) is trained over 100 epochs using sparse categorical cross-entropy loss and the Adam optimizer, achieving 93.4% training and 86.1% held-out evaluation accuracy. Generation employs temperature sampling (T = 1.2) with a stochastic seed-mutation strategy (μ = 0.10). Candidates are ranked by a six-criteria physicochemical scoring function grounded in the Kyte–Doolittle hydrophobicity scale and Boman charge index, and top candidates are evaluated by ESMFold v1, providing per-residue pLDDT confidence scores. Results yield quality scores of 6.17–6.48, novelty similarity values of 0.12–0.19 (below the 30% identity threshold accepted as the community standard), and mean pLDDT scores of 68.4–72.1. With only 78,356 trainable parameters and under 30 minutes of training on a single consumer GPU, the framework matches or surpasses the novelty and validity rates of transformer-scale systems while remaining accessible to resource-constrained academic laboratories.
Objective:
Designing functional proteins from first principles confronts a combinatorial barrier of staggering scale: a polypeptide of 100 residues spans 20^100 possible sequences, a search space entirely inaccessible by experimental means alone. This paper presents an end-to-end artificial intelligence pipeline that learns the statistical grammar of proteins directly from the UniProtKB Human Proteome database and generates novel, biologically plausible candidates without structural templates or homologous scaffolds. A character-level Gated Recurrent Unit (GRU) model (64-dimensional embedding, 128 hidden units, 78,356 parameters) is trained over 100 epochs using sparse categorical cross-entropy loss and the Adam optimizer, achieving 93.4% training and 86.1% held-out evaluation accuracy. Generation employs temperature sampling (T = 1.2) with a stochastic seed-mutation strategy (μ = 0.10). Candidates are ranked by a six-criteria physicochemical scoring function grounded in the Kyte–Doolittle hydrophobicity scale and Boman charge index, and top candidates are evaluated by ESMFold v1, providing per-residue pLDDT confidence scores. Results yield quality scores of 6.17–6.48, novelty similarity values of 0.12–0.19 (below the 30% identity threshold accepted as the community standard), and mean pLDDT scores of 68.4–72.1. With only 78,356 trainable parameters and under 30 minutes of training on a single consumer GPU, the framework matches or surpasses the novelty and validity rates of transformer-scale systems while remaining accessible to resource-constrained academic laboratories.
Methods:
Protein Design Using Gated Recurrent Units
Results:
The trained GRU model generated protein sequences of 130 amino acids, all within the 80–200 AA optimal window. Table 1 presents the top-three sequences, including quantitative ESMFold pLDDT structural confidence scores. Quality scores cluster between 6.17 and 6.48, consistently above the 5.0 plausibility threshold. Novelty similarity values of 0.125–0.188 confirm genuine novelty by the community-accepted standard. Mean pLDDT scores of 68.4–72.1 indicate confident local structure prediction—a structural quality signal entirely independent of the sequence-level scoring function.
Conclusions:
This paper has presented an end-to-end, resource-accessible framework for de novo protein sequence generation using Gated Recurrent Units. By tightly integrating GRU-based character-level language modeling, a six-criteria literature-grounded physicochemical scoring function, stochastic novelty verification, and ESMFold real-time structure prediction with quantitative per-residue pLDDT confidence reporting within a single unified pipeline, the system enables structurally grounded protein design on hardware available to any research group with free cloud GPU access.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.