Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Bioinformatics and Biotechnology (no longer under consideration since Feb 19, 2022)

Date Submitted: Jan 5, 2022

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Identification of somatic mitochondrial variants in colon cancer patients by in-silico analysis and next generation sequencing

  • Gadicherla Ramya; 
  • Venkateswara Rao Talluri; 
  • Niraj Rai; 
  • Smita C Pawar; 
  • Rajath Othayoth

ABSTRACT

Background:

Cancer is a result of mutations in the genome. Apart from the ~3.2GB of nuclear genomic DNA, human cells have hundreds to thousands of mitochondria, each of which carry few copies of circular mitochondrial genome 16,569bp [1]. Mitochondria are essential organelles in eukaryotes that perform key functions ranging from the generation of bioenergetic intermediates like ATP and GTP to Nucleotide synthesis, FE-S clusters, Haem, amino acids, Fe2+/Ca2+ handling, inflammation and apoptosis [2].The positional virtue at such a cellular cynosure/epicenter, mitochondrial dysfunction and consequential metabolic defects are reported in various human pathologies that include different forms of sporadic and familial cancers [3]. Mammalian cells under normal conditions rely on Oxidative phosphorylation for generation of major supply of ATP. In contrast, cancer cells show increased ATP production through non-oxidative process of breakdown of glucose in cytoplasm [Warburg Effect, 4]. This metabolic reprogramming is considered as a hallmark of cancer [5]. However, majority of ATP in tumors still show dependence on oxidative phosphorylation for proliferation [6]. Evidently all tumors arise as a consequence of acquired genetic alterations [7]. The studies of tumor-associated mutations are facilitated by massive parallel sequencing which has provided major insights into the changes occurring in nuclear genome [8, 9]. However, in spite of the fact that mitochondrial dysfunction is a characteristic property of any tumor cell, few studies have utilized the high throughput sequencing methodologies to characterize the mitochondrial genome mutations [10, 11]. The mitochondrial DNA is a double-stranded circular molecule of 16,569 base pairs that harbors 37 genes of which 22 genes code for transfer RNA’s, 2 genes code for ribosomal RNA and 13 genes code for proteins that are involved in Oxidative phosphorylation [OXPHOS] process [12]. A striking distinctive feature remained with an asymmetry between the two strands of mitochondrial DNA in terms of gene distribution and nucleotide content [13]. The two strands include ‘G’-rich Heavy strand[H-Strand] which is a template for transcription of most of the proteins(12 out of 13) and a ‘C’-rich Light strand[L-Strand] that has one protein coding gene[ND6 ] that is transcribed. Anecdotal studies have indicated the presence of mitochondrial DNA mutations in cancer literature for several decades [14,15] yet mitochondrial genetics has been largely neglected. With the current availability of large datasets and better analytical approaches, the presence of at least one mitochondrial DNA mutation in approx. 60% of solid tumors was demonstrated that do not fit in patterns associated with oxidative damage [16, 17, 18]. Mitochondrial DNA is more prone to mutations when compared to nuclear DNA due to lack of intronic regions, DNA repair machinery and histone proteins. Most of the variants are likely to reside within the coding regions of mitochondrial DNA [19] as the non-coding region accounts for only ~38% of the entire mitochondrial genome [20]. Conversely in the disease state, mutations are observed throughout the mitochondrial genome [21, 22]. Somatic variations are of major significance pertaining to cancer. These variations may cause amino acid alterations in the resultant protein product and may exert effects that may be deleterious on the structure, function, solubility and stability of proteins.[23].Hence, it is possible that somatic variations play a major role in diversity [functional] of proteins and may often have association with human diseases. Studies have shown the association of somatic variations to many disorders and cancer [24]. Examining the effect an amino acid variant on the function of the encoded protein using various in-silico SNP detection tools can enhance the choice of a functionally polymorphic Somatic variation in an association study. Deleterious or disease-causing mutations are often observed at the core or interface residues that mediate protein interactions. The core residue mutations may destabilize protein structure whereas interface residue mutations can affect the binding energies of protein-protein interactions. However within interfaces, small subsets of residues are essential for binding the free energy of the protein–protein complex [25]. Protein-protein interactions are perilous for almost every process in the cell and mutations that hinder these interactions can have dire consequences on the cellular function associated. The interface mutations are often considered as important contributors to human disease. The impact of mutations on the native structure of a protein can help in understanding the stability, solubility and protein-protein binding affinity. Association of mitochondrial DNA variants in colorectal cancer (CRC) patients has been studied in various populations [26]. CRC is regarded as the third most common cancer worldwide and is the second leading cause of cancer-related mortality [27]. In India, CRC is the fourth most prevalent in men and third most in women [28]. Risk factors include environmental factors such as lifestyle and diet [29]. Genetic basis of CRC has also been established in various studies [30,31]. CRC carcinogenesis is a complex process usually beginning as a benign growth called polyp on the inner wall of colon or rectum. The polyp may remain benign or transform into adenomas, which have the potential to transform into cancer over a period of time under the influence of factors that may also be genetic[32].CRC genomes, like any other solid tumors, harbor genetic alterations that range from small-scale deletions, insertions, point mutations to large scale copy number changes and rearrangements. Such type of alterations may confer to colorectal carcinogenesis. The alterations may be found in nuclear as well as mitochondrial genome. Significant progress has been made in characterizing the process of CRC carcinogenesis at epigenetic and genetic level [33]. A study was conducted by Cheng-Ye Wang et.al.2011, on Colorectal cancer patients in Chinese population by sequencing the complete mitochondrial genome of tumor and near normal tissue from the same patient using Next Generation Sequencing. Findings revealed identification of somatic variants in both tumor and normal tissues though at differing frequencies suggesting altered physiological environment in cancer patients. [34] In the present study, a similar variants identification approach unique to Colorectal cancer with a focus on somatic variants in the coding regions of mitochondrial genome was adopted using NGS in selected population of Andhra Pradesh and Telangana. As the number of mutations identified were very high, it may not be possible to validate all of them functionally by trans mitochondrial studies. In silico method helps to narrow down initially on the potentially pathogenic variants, which may pave a way to functionally validate the reported mitochondrial DNA variants in in vivo/in vitro model systems. Various In silico techniques were applied for screening, filtering and analyzing unique variants [35]. The analyses tools include SIFT, Polyphen 2, Mutation Assessor for pathogenicity prediction and SNP&Go for Protein stability prediction. Homology modeling was used to study protein 3-dimensional structural impacts on both native and mutant proteins and their effects viz., stability, solvent accessibility and protein-protein interactions.

Objective:

Mitochondrial DNA variants have been shown to play a key role in colorectal cancer progression. In the present study, some of the somatic variants causing dysfunctioning of OXPHOS pathway leading to carcinogenesis were identified. These variants can play a vital role as biomarkers and drug targets as well. In the current study, somatic variants were identified by sequencing mitochondrial genome of tumor and near normal colorectal tissue using Next Generation Sequencing (NGS) technology. Around 14 variants were found across majority of samples in mitochondrial genes of Complex-I, Complex-IV and Complex-V of tumor tissues. Initial characterization of the 14 variants were carried out using insilico predictions using tools like SIFT, Polyphen2, Mutation Assessor, and SNPs & GO for pathogenicity. 3D structural analysis was also carried out by homology modeling of the mitochondrial proteins for both wild type and mutated proteins. Variant S149T [T6348A] in MT-CO1 gene has been predicted as a deleterious mutation across all the insilico tools. The mutation was found to increase flexibility of the MT-CO1 protein structure impacting the protein functioning. A stop-gain mutation at position T9038G of MT-ATP6 gene may lead to truncation of the protein which may further influence activity of the whole MT-ATP6 protein across the complex. Thus, the variants identified in this study may play a key role in colon carcinogenesis as disruption in function of mitochondrial genes can influence the OXPHOS process.

Methods:

Study population and ethics statement This study comprised of consecutive patients with CRC from Telangana and Andhra Pradesh and was approved and conducted under the surveillance of the institutional human ethics committee of MNJ Institute of Oncology, Hyderabad. Patients who were diagnosed specifically for CRC were included and patients with other disorders like inflammatory bowel disease [IBD], familial adenomatous polyposis [FAP], Lynch syndrome, patients undergoing cancer therapy were excluded from the study. Sample Collection and DNA Isolation Tumorous tissue along with an adjacent normal tissue within surgical safety border (about 5 cm from the tumor) were excised and immediately transferred to 4ml of 96-100% Ethanol and frozen at –20 ̊C until further analysis. The stage of the cancer was classified as T1, T2, T3 and T4 respectively [Histopathological analysis]. A total of 25 samples [25 Tumor and 25 near normal] were included in the study. Tissue samples (50 mg) were immersed in 3mL of TE buffer [pH-8.0] and incubated at 37°C for 2hrs with intermittent buffer changes. The buffer change was repeated thrice every 2 hrs. After the third buffer change, the tissue was suspended in 2ml TE buffer, incubated overnight at 37°C for the removal of any ethanol residues. The tissues were then homogenized and DNA was extracted using total DNA isolation kit [Bioserve India Pvt. ltd., India.]. Qualitative and quantitative assessment of DNA was carried out by Agarose Gel electrophoresis and UV spectrophotometric analysis respectively. Further quantitative assessment was done by Qubit fluorometry to measure the accurate amount of double stranded intact DNA. PCR amplification of mitochondrial DNA Whole mitochondrial DNA amplification (two fragments of ~8.5 kb each) was carried out using a long range PCR. The primer sequences to amplify first amplicon were AAATCTTACCCCGCCTGTTT (forward), AATTAGGCTGTGGGTGGTTG (reverse) and for second amplicon GCCATACTAGTCTTTGCCGC (forward) and GGCAGGTCAATTTCACTGGT (reverse) [36]. The long range amplification was done by using TaKaRa LA Taq® DNA polymerase kit [Clontech, South America] as per the manufacturer’s protocol. PCR conditions cast off were 94°C for 1 min, 30 cycles of 94° for 30 seconds, 54°C for 30 seconds, 72°C for 9 minutes with a final extension of 72°C for 5 minutes. After amplification, the products were run on a 0.8% agarose gel and expected fragments of 8.5kb were excised. PCR products were gel purified using gel extraction kit [Bioserve India Pvt. Ltd.]. Library preparation and sequencing The purified amplicons [50ng-100ng] were taken up for DNA library preparation. All the samples were quantified by Qubit Flouorometry using dsDNA HS Assay Kits [Thermofisher Scientific, USA] on Qubit 2.0 Fluorometer[Thermofisher Scientific, USA]. Ion Shear Plus Kit was used for enzymatic fragmentation of the amplicons to size of 200bp. The fragmented library was further subjected to ligation and final library amplification using NEBNext Fast DNA Library Prep Set for Ion Torrent [NEB] and Ion Xpress Barcode Adaptors kits (Ion torrent, Thermofisher Scientific, USA]. Intermittent purifications and size selection [200bp libraries] was carried out using Agencourt AMPure XP magnetic bead based purficiation system (Beckman Coulter,USA) according to manufacturers guidelines (NEB,USA). The library profile and quantification was done on Bioanalyzer 2100 System (Agilent Technologies, Santa Clara, CA, USA) using high sensitivity DNA analysis kit. Libraries thus quantified were taken as template for Clonal amplification and further for sequencing on Ion Proton system. Mitochondrial DNA sequencing analysis The Ion Proton sequencing data was further taken up for the analysis using CLC Genomics workbench version 20.0.4 [QIAGEN]. The Raw sequencing data was mapped against the Revised Cambridge Reference Sequence (rCRS) (Andrews et al. 1999). Quality parameters include the exclusion of polyclonal reads and Phred scores lower than 20[>=Q20]. Variants were called using the option of Low frequency variant calling workflow on CLC Genomics workbench. The parameters selected for variant calling include minimum coverage of 10, mismatch of 2, and minimum quality of base 20. The minimum base coverage was taken as 100. A minimum threshold of 2% was applied for SNP’s. The sequences were checked for any possible contamination using reported mitochondrial DNA polymorphisms. As the mitochondrial DNA is small [around 16569bp] and explored extensively, most of the mitochondrial DNA germ line polymorphisms were already known. Hence, when a tumor sample is contaminated by other samples, the somatic-like mitochondrial DNA substitutions by contaminants were likely to be overlapped with known mitochondrial DNA polymorphisms. Therefore, the variants were further filtered basing on the information available about the variants in mtDB database and Mitomap for uniqueness. The mutations, which were not reported in Mitomap, were considered as rare and taken up for further analysis. Copy number Analysis by RT-PCR The relative abundance of mitochondrial DNA [Mitochondrial copy number] was assessed by real-time quantitative PCR. In the current study, the samples ND1 specific for mitochondrial DNA and beta-actin specific for Nuclear DNA were selected. The sequence of primers used for amplification of ND1 was: forward primer:’5’-TAATGCTTACCGAACGAA-3’ Reverse primer:’5’-TTATGGCGTCAGCGAAGG-3’ which amplify a 104bp fragment. The sequence of primers used for amplification of 105bp fragment of beta-actin are: forward primer-5‘-GCAAAGTTCCCAAGCACA-3’Reverse Primer-5’AAGCAAGCAGCGGAGCAG-3’.[37]The PCR cycling conditions used were: 50 for 2mins, 95°C for 10mins, followed by 35cycles of 95 for 15sec , 50 for 30sec. Melt curve or dissociation curve is also calculated to ensure presence of single PCR product. Each sample was setup in triplicates. For each clinical sample, the Ct values for mitochondrial DNA and nDNA were determined. Average of Ct values were taken and further calculated the copy number by using the following equations [38]: ΔCt= (nucDNACt-mitochondrial DNACt) Relative mitochondrial DNA content =2 X 2ΔCt . In silico analysis for pathogenicity prediction One way to identify the pathogenicity of a variant or mutation is to study the biochemical effect it exhibits. As it is not possible to study effect individually for every mutation, computer based algorithms could be used to evaluate the effects from substitutions. In silico based approaches including various tools such as SIFT, Polyphen 2, Mutation Assessor and SNPs & GO were used in this study. a) SIFT (https://sift.bii.a-star.edu.sg/) and Mutation Assessor (http://mutationassessor.org/r3/) are sequence homology based approaches to identify positions that are evolutionarily conserved between and within species using multiple alignments of distantly related and homologues sequences. Conservation of amino acid substitution across multiple sequences will be presumed important for function of a protein. b) Polyphen2[http://genetics.bwh.harvard.edu/pph2/] uses sequence homology together with a combination of extracted data, such as protein function annotations from Uniprot database and secondary structure information from DSSP program crystallographic D-factors from RCSB protein databank[PDB] and calculated structural parameters including accessible surface area, hydrophobicity and electrostatistics to arrive at a prediction. IF the alignments, annotations or basic structural parameters suggest the conserved substitution buried in the protein structure within an active site or part of disulfide bridge or protein-protein interaction region, it will be predicted pathogenic. c) SNPs & GO (http://snps.biofold.org/snps-and-go/snps-and- go.html) are support vector machine [SVM] based classifiers which takes the input protein sequence profile and functional information and predict the effect of non-synomymous variation. Inspite of their predictive potential, these algorithms can yield contradictory information, one of the reasons being the lack of human mitochondrial protein structures. These programs utilize sequences of lower organisms for analysis. Hence, 3D structural analysis was employed to determine the functional importance of amino acid substitutions in mitochondrial proteins. 3D-Structural Analysis of Mitochondrial Proteins Stability of native and mutant protein models were performed by structural analysis. The three dimensional [3D] structures for the proteins coded by the genes were identified using BLAST web resource. Homology modeling approach was used for prediction of 3D structure for the proteins whose structures were not available. The homology modeling was performed using program Modeller-10.1 (https://salilab.org/modeller/). The 3D models for the proteins MT-ND1, MT-ND2, MT-CO1,MT-CO2,MT-ATP6,MT-ND4 and MT-ND5 were constructed using MODELLER 10.1 software. Atleast 5 models for each of native and mutant protein were created for the proteins under study. Each model was validated using PROCHECK,[https://www.ebi.ac.uk/thornton-srv/software/PROCHECK/]stereo chemical quality was assessed by RAMPAGE[http://mordred.bioc.cam.ac.uk /~rapper/rampage.php] and ProSA server and the best model was picked for further analysis. Energy minimization was carried out using Chiron server. The divergence and deviation between the native and mutant structures was evaluated by RMSD values using SUPERPOSE[http://superpose.wishartlab.com/]. UCSF chimera [ver. 1.15, http://www.cgl.ucsf.edu/chimera/] software was used to check the changes in hydrogen bonding, clashes and contacts for each residue upon mutation. The following tools were further employed to assess the impact of mutations on various dynamics of the proteins such as protein-protein interactions, solvent accessibility, protein stability, free-energy changes and so on. a) InterProSurf -Predicts protein-protein interaction effects. It predicts interacting residues on a protein subunit and locates the interface residues in a protein complex. b) HOPE web server was used to study the structural effects of mutations with an actionable hypothesis. The predictions combine the structural information from several sources, including the 3D protein structure, sequence annotations [Uniprot] and Distributed Annotation System (DAS) server predictions. All the collected information is used to analyze the effect of the mutation on the protein structure. c) DynaMut-[http://biosig.unimelb.edu.au/dynamut/related] a web server that can be used to analyze and visualize the dynamics of a protein by sampling conformations and impact assessment on stability that is a result of vibrational entropy changes.

Results:

Subject characteristics The present study was carried out to identify mitochondrial somatic variants in mitochondrial genome of CRC tissue samples. Demographic data was presented in Table-1. Copy Number Analysis Relative copy number analysis was calculated for all the samples using nuclear gene Beta-actin and mitochondrial gene ND1. The details of the copy number variations in comparison to adjacent normal tissues have been tabulated [Supporting Information-A1]. Sequence analysis of mitochondrial DNA Sequencing of mitochondrial genome was carried out on a total of 50 samples [25 Tumor and 25near normal] of which, 15 samples were of Colon cancer and 10 of them were of Rectal cancer. Quality check of the raw reads was carried out and low quality reads with Phred scores lower than 20[< = Q20] and polyclonal reads were filtered out. The refined raw data was then mapped against the Revised Cambridge Reference Sequence (rCRS) (Andrews et al. 1999) with coverage of approximately 300x for each sample was achieved. Variants were filtered by applying the parameters of quality and threshold. In order to minimize any errors, variants that were supported by at least five reads were considered in the study. The study concentrated more on the coding regions of mitochondrial genome and hence the variants of D-loop and RNA’s were excluded from this study. A total variants obtained across the samples ranged between 220-310. Most of the variants across colon samples were observed in MT-ND5 [20%] gene followed by MT-CYTB [13%], MT-ND1[9%], MT-ATP6[8%] and MT-ND4[7%] . Variants were observed in more number in MT-ND5 [22%] gene followed by MT-CYTB [15%], MT-ATP6 [7%], MT-ND2[6%] and MT-ND6[2%] in rectal cancer samples. Around 18 variants were selected across 25 cancer tissue samples that were not found in the matched normal samples[Supplementary Information-A2].These variants span the mitochondrial coding regions of Complex-I, Complex-IV, in Complex-V. All the variant positions were checked for any possible contamination or their origin as germ line variants by submitting to mtDB database as the database is repository for most of the known germ line polymorphisms. When a tumor sample is contaminated by other samples, the somatic-like substitutions are likely to overlap with known mitochondrial DNA Polymorphisms. Hence, variants at positions 4122[A>C], 4172[T>A], 4820[G>C] and 7985[C>G] were excluded from this study. The remaining variants were checked for their pathogenicity using Insilco tools to determine the biochemical effect of the site of mutation [Table-2]. Structural Analysis of Proteins Homology modeling method helps in mapping variants to protein structure. The 3D structures of proteins under study such as MT-ND1, MT-ND2,MT-CO1, MT-CO2, MT-ATP6 MT-ND4 and MT-ND5 were created using the Modeller 10.1 software and the structural homologs used as a reference were tabulated . Energy minimization of the modeled protein structures using Chiron tool yielded final energy after minimization and the RMSD for native as well as mutant protein[Supplementary Information-A3]. The structures on validation using Z scores by PROCHECK, RAMPAGE and ProSA[z-score] predicted overall model quality and the total energy deviation of the structure using random conformations. The modeled structure was anticipated to be inaccurate if the z-scores range outside the characteristic for native proteins. The variants F242L, W290R, P79A, I19N,Y174S &S461G exhibited increase in the total energy values whereas the other variants showed decrease in the total energy values. Further Interprosurf tool depicts the number of surface atoms and buried atoms in a protein structure. The variants L150P [T3755C], F242L[C4032G], W290R [T4174C],L249P[T5215],F260V[T5247G], S149T [T6348A], P79A[C7820G],L140P [T12755C],Y174S [A12857C]& S461G [A13717G] were found to have less number of buried atoms when compared to the respective wild type. No change was observed in variant position W68S[G8729C]. For positions S339R [A11774C]& I19N [T12392A], number of residues were found to be buried [Supplementary Information-A4] The actionable hypothesis of each of the variants was predicted using HOPE server. The MT-ND1 gene variants T3755C[L150P], C4032G [F242L] were predicted not to damage the protein and T4174C[W290R] possibly damaging variant based on the conservation score. The variant on MT-ND2 gene T5215C[L249P] &T5247G[F260V] may be probably damaging. MT-ND4 variant A11774C[S339R] may hamper the protein activity. Variants T12392A[I19N], T12755C [L140P], A12857C[Y174S]on MT-ND5 gene might be damaging whereas A13717G[S461G] might disrupt the flexibility. The variant T6348A[S149T] on MT-CO1 gene and MT-CO2 gene variant C7820G[P79A] may not be damaging but may disrupt the activity of protein. Mutation on MT-ATP6 gene G8729C[W68S] may disrupt the function of protein. DynaMut tool was used to assess the stability changes of a protein upon acquiring a mutation. Of the 13 variants, T4174C [W290R],T3755C[L150P],C7820G[P79A],G8729C[W68S]&A11774C[S339R] were found to be stabilizing the structure of the protein. The remaining 8 variants were found to have destabilizing effect. The vibrational energy entropy between wild type and mutant was found to increase the molecular flexibility[Table: 3]. Interatomic changes which are essential for point mutation studies were carried out within wildtype and mutant that allows the visualization of interactions. Changes across all the mutants showing bond disruptions were depicted in Figure 1. Deformation energies and Molecular and atomic fluctuations were also carried between wild type and mutant protein PDB structures were depicted in Figure 2. Discussion The present study identifies somatic mitochondrial mutations in colorectal cancer from cancerous as well as near normal tissues that can be used as markers in the prognosis and diagnosis. Mitochondria are energy centers of a cell that produce ATP via oxidative phosphorylation [39] and an important role in apoptosis [40-45] though Otto Warburg published a hypothesis that reported production of ATP through glycolysis indicating the potential role of defective mitochondria in carcinogenesis. Similar studies carried out on 2854 patients[UK Population] with single nucleotide polymorphisms (SNPs) [46] and 2838 patients originating in Scotland [47] indicated the pathological effect of detected mutation and its impact on cell phenotype. Studies have found somatic mitochondrial mutations and copy number variations in colorectal tumors suggest a significant role in carcinogenesis. [48,49]. A relative copy number analysis indicates variation between samples of same stages. The copy number was found to be increased in four samples of stage-I, five samples of stage-II and six samples each in stage-III and IV. In contrast, the copy number was found to decrease in two samples each of stage II& III. Possible explanation for high copy numbers might be good oxygenation during the stage progression. Low copy number may be due to lower oxidative stress that is advantageous for cancer progression. The parameter can hence be utilized to monitor the effects of therapy in the diseased condition. In the present study, samples were subjected to mitochondrial genome sequencing by NGS technology on Ion torrent platform. Applying various criteria such as read quality read coverage, frequency of occurrence and so on filtered variants. Nevertheless, these variants can be used for screening purposes as markers for prognosis and diagnosis, as these were observed across majority of samples in this study. A total of 10 variants were detected spanning Complex-I genes [such as MT-ND1, MT-ND2, MT-ND4 & MT-ND5] whereas 2 variants detected in Complex-IV [MT-CO1&MT-CO2] and 2 variants detected in Complex-V [MT-ATP6] and analyzed using various insilico tools. Only one variant T6348A has been predicted pathogenic across all the tools and results varied for others. An actionable hypothesis was derived by submitting the variants to HOPE server that makes use of information from various sources that may be useful for predicting the effects of variants. The variants on MT-ND1 gene T3755C [L150P], C4032G [F242L] were found to be neutral, whereas the variant T4174C [W290R] was possibly damaging as the wild type amino acid was found to be conserved. At all the three positions, the size of mutant amino acid was smaller than the wild type, though not damaging but can affect the contacts with lipid-membrane. The variant on MT-ND2 gene T5215C [L249P] might disrupt the function of protein, as it was located in the main core of protein responsible for activity. Also mutation T5247G [F260V] may be probably damaging as it was located in highly conserved domain. MT-ND4 variant A11774C [S339R] was located in highly conserved position and domain required for activity and hence disruptions are possible leading to loss of protein activity due to mutation. Variant T12392A [I19N] on MT-ND5 gene might be damaging as the size of mutant residues were bigger than the wild type residue and located in the main core domain, which may lead to disruption of protein function. Variant positions T12755C [L140P] & A12857C [Y174S] might be damaging as they were located in the core domain of the protein. The variant A13717G [S461G] introduces Glycine, which was very flexible and can disrupt the rigidity required by the protein at this position. The variant position T6348A [S149T] on MT-CO1 gene shows that the size of the mutated moiety was larger than the wild type and mutated residue in located in a domain that was important for protein activity and in contact with another domain also important for the activity. The mutation may disturb the interaction between these domains, which might affect the function of protein. The variant C7820G [P79A] on MT-CO2 gene may disturb the protein confirmation as Proline, the wild type entity was known to be rigid conferring a special backbone confirmation, which might be disrupted due to mutation. Also, mutated residue was located in the transmembrane domain important for activity of protein and hence may lead to disturbance in interaction with other domains. Mutation on MT-ATP6 gene G8729C [W68S] may disrupt the function of protein, as it was located in the domain important for main activity of the protein. Hydrophobicity differences between wild type and mutant residues lead to loss of hydrophobic interactions. Despite the potential to predict the effects of amino acid substitutions, the tools can perform sub optimally and yield contradictory results. One of the reasons being lack of human mitochondrial protein structures of the complexes. The programs used earlier utilize sequences from lower organisms for analysis. The role of specific amino acid substitution can be misinterpreted or rather missed given the structural, functional differences between distantly related templates. Though, the main features of catalytic subunits are conserved widely, it gets complicated when dealing with eukaryotic cells especially where new interfaces are generated from subunits. Hence, 3D structural analysis was utilized for bigenomic, quaternary structural organization of mitochondrial proteins and to determine importance of amino acid substitution at functional level. Regardless of the role of protein elements that are fundamental to its activity, computational recreation has played a major part to decipher the impact of mutations on the structure and function of a protein by relying on their static structures. The 3D structures were constructed using Modeller 10.1 tool. Initially native models of all the mitochondrial proteins were constructed. Further, mutations were introduced into the protein sequence and mutant models were constructed. All the constructed models were evaluated using PROCHECK, RAMPAGE and ProSA to calculate any deviations between the native and mutant structures. Dynamut server can assess the structure of protein and evaluate the role of mutations on strength of protein based on changes in vibrational energy and dynamicity. A cut off set for this program is ΔΔG<0 is destabilizing and ΔΔG≥0 will be a stabilizing one as summarized in Table: 3. The vibrational entropy energy changes between wild type and mutant was also calculated. For the residues L150P, W290R, W68S & S339R the molecular flexibility decreased. Mutations of F242L, L249P, F260V, S149T, P79A, I19N, L140P, Y174S & S461G increase molecular flexibility due to interatomic changes, disruption of hydrogen bonds, halogen bonds and hydrophobic interactions between different atoms in a protein. Overall, the mutations at the specified sites of the genes will eventually cause disruption in the normal oxidative phosphorylation process. Earlier reports suggest that ND subunits of complex-I play an important role and mutations in these subunits impair the oxidative phosphorylation and increase oxidative stress within the cell. Both compromised and induced function due to accumulation of mutations of Complex-IV and Complex-V genes have been observed in colon carcinoma when compared to normal mucosa according to earlier studies. [50] The present in silico studies indicate functional impact of mutations in mitochondrial genes that could lead to colorectal cancer tumors. The 3D structural analysis led to analyze the key role played by mutations and functional importance of amino acid substitutions in cancer. Though, there have been discrepancies observed across mutational impact across various tools, it may give an overall idea as to which mutations can be taken up for further studies.

Conclusions:

Impact of amino acid substitutions on stability of a protein remains one of the most important setbacks in protein science. A new approach including the experimental and computational data analysis has an added advantage. In this study, somatic mutations were identified using high end NGS sequencing technology and further analysis were done by in silico methods. The 3D structural analysis plays a major role in identification and filtering of mutations that were functionally impacting the proteins. In the present study, a single mutation S149T in MT-CO1 gene was identified that has yielded similar analysis outcome across all the tools. Further experimentation and screening of across more number of samples is essential to conclude any specific mutation in mitochondria of colorectal cancer in the present population. Clinical Trial: Not applicable


 Citation

Please cite as:

Ramya G, Talluri VR, Rai N, Pawar SC, Othayoth R

Identification of somatic mitochondrial variants in colon cancer patients by in-silico analysis and next generation sequencing

JMIR Preprints. 05/01/2022:35952

DOI: 10.2196/preprints.35952

URL: https://preprints.jmir.org/preprint/35952

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.