Font Size: a A A

Landscape And Hotspots Of Spontaneous Mutations In The Bacterium Escherichia Coli By High-throughput Sequencing

Posted on:2019-03-13Degree:MasterType:Thesis
Country:ChinaCandidate:X L ZhangFull Text:PDF
GTID:2394330545493469Subject:Cell biology
Abstract/Summary:
Objective : Spontaneous mutation is the source of driving force of biological evolution.The investigation of spontaneous mutation by Next-Generation Sequencing technology has attracted a large amount of attention recently due to the fundamental roles of spontaneous mutation and pathological process.However,most of these studies are based on the Mutation Accumulation Experiment and the Long-term Evolution Experiment.The majority of mutations are accumulated by long-term culture of bacteria strains,which will hinder the investigation of the distribution of spontaneous mutation in the whole genome.In this study,we developed a molecularly-barcoded deep sequencing strategy and binomial distribution algorithm to obtain high confident spontaneous mutations(P<0.01),which could enable us to investigate the distribution of spontaneous mutation hotspot in the whole genome.Method: The Duplex Sequencing(DS)dataset of Tp53 exon4-10 and Improved Duplex Sequencing(IDS)dataset of E.coli whole genome are used in this study,and both of the dataset have 15 samples and sequenced on Illumina HigSeq 2500 platform with paired-end reads of 150 bases.The dataset of Tp53 exon4-10 was firstly processed to calculate the 12 error types of sequencing error by grouped the reads,origin from same DNA fragment,into a Family Group and construct a consensus sequence.We developed Improved Duplex Sequencing(IDS)technology,which uses four kinds length of random bases as molecular tags to improve base balance and sequencing quality.The IDS technology was used to construct library for 15 samples of E.coli whole genome,and binomial distribution method was performed to obtain high confidence(P<0.01)spontaneous mutation.Finally,we use the RLE(Run-Length Encoding)algorithm to extract hotspot regions formed by spontaneous mutation in the genome wide and annotate the genes in the hotspot regions shared by the 15 samples to study the distribution of mutation in these genes.Results: By processing the DS dataset of Tp53 exon4-10,we found that the error rate of NGS technology is very close to 0.001 provided by Illumina,but the difference among all error types are notably,which indicate that different error types should be considered separately.Through the improvement of original Duplex Sequencing strategy,we can solve the problem of conventional NGS in removing PCR duplication effectively.In addition,this sequencing method can also improve the base balance and quality of sequencing data.We obtained lots of high confident(P<0.01)spontaneous mutation by binomial distribution method,and RLE algorithm was performed to obtain 3 notably spontaneous mutation hotspots,which were validated in all 15 samples.Finally,we annotated the genes in these three hotspots and found that most of the mutations were found in the repetitive or non-functional regions of the gene.Conclusion: In this project,we found that conventional NGS technology is not suitable for high-throughput sequencing of short gene sequence.In addition,the consistency of the hotspots in the 15 samples indicate the effectiveness of mutation filtering method based on binomial distribution.The distribution of spontaneous mutations in the hotspot genes can also provide a basic evidence for the investigation of biological evolution.
Keywords/Search Tags:Escherichia coli, Spontaneous mutation, Deep Sequencing, Binomial distribution
Related items