Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Jan 17, 2024)
Date Submitted: Oct 4, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Multiomics Learning to Extrapolate Proteome Expression: Deep learning approach with fast inference validated for various human cancer types
ABSTRACT
Background:
Complex multiomics data requires interpretation on different levels of molecular biology processes. However, genomics and transcriptomics methods are more sensitive than proteomic methods because of PCR reaction and ability to multiply molecules to level enough for a successful detection. This forms the gap between thousands of genes with known RNA expression and hundreds of detected proteins.
Objective:
The solution was designed as an extrapolating tool for online public web-usage, because of that only light-weighted deep neural architectures with fast inference were used.
Methods:
To predict protein abundance by known RNA level in a sample we have used a multi-channels input deep convolutional neural network to store label-free LC-MS/MS data, amino acids sequence and database annotations in one tensor. Gene mappings databases with many-to-many relationships were used to construct input feature maps.
Results:
We have considered two datasets with normal human tissues and human tumor cell lines. Healthy human tissues dataset Tissue29 contained 12417 proteins for 29 samples which leads to 251349 experiments. Human tumor cell lines dataset NCI60 contained 251706 experiments on 46 samples for 7812 proteins. The validation set was once created by randomly selected 20% of proteins in both of used datasets. Validation results for the healthy human tissues dataset showed 0.58 R2 coefficient of determination and 0.71 Spearman correlation for all samples and up to 0.67 R2 per tissue. For the tumor cell lines dataset 0.54 R2 coefficient of determination and 0.71 Spearman correlation for all validation samples and up to 0.62 R2 per cell line.
Conclusions:
The solution allows a researcher to use any annotation collections and modifiable amino acids sequence to extrapolate protein abundances with known RNA quantity in the sample. A “one model – many datasets” pipeline with comparable metrics was implemented.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.