Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
285 result(s) for "genotype imputation"
Sort by:
GeneImp: Fast Imputation to Large Reference Panels Using Genotype Likelihoods from Ultralow Coverage Sequencing
We address the task of genotype imputation to a dense reference panel given genotype likelihoods computed from ultralow coverage sequencing as inputs. In this setting, the data have a high-level of missingness or uncertainty, and are thus more amenable to a probabilistic representation. Most existing imputation algorithms are not well suited for this situation, as they rely on prephasing for computational efficiency, and, without definite genotype calls, the prephasing task becomes computationally expensive. We describe GeneImp, a program for genotype imputation that does not require prephasing and is computationally tractable for whole-genome imputation. GeneImp does not explicitly model recombination, instead it capitalizes on the existence of large reference panels—comprising thousands of reference haplotypes—and assumes that the reference haplotypes can adequately represent the target haplotypes over short regions unaltered. We validate GeneImp based on data from ultralow coverage sequencing (0.5×), and compare its performance to the most recent version of BEAGLE that can perform this task. We show that GeneImp achieves imputation quality very close to that of BEAGLE, using one to two orders of magnitude less time, without an increase in memory complexity. Therefore, GeneImp is the first practical choice for whole-genome imputation to a dense reference panel when prephasing cannot be applied, for instance, in datasets produced via ultralow coverage sequencing. A related future application for GeneImp is whole-genome imputation based on the off-target reads from deep whole-exome sequencing.
I am hiQ—a novel pair of accuracy indices for imputed genotypes
Abstract Background Imputation of untyped markers is a standard tool in genome-wide association studies to close the gap between directly genotyped and other known DNA variants. However, high accuracy with which genotypes are imputed is fundamental. Several accuracy measures have been proposed and some are implemented in imputation software, unfortunately diversely across platforms. In the present paper, we introduce Iam hiQ, an independent pair of accuracy measures that can be applied to dosage files, the output of all imputation software. Iam (imputation accuracy measure) quantifies the average amount of individual-specific versus population-specific genotype information in a linear manner. hiQ (heterogeneity in quantities of dosages) addresses the inter-individual heterogeneity between dosages of a marker across the sample at hand. Results Applying both measures to a large case–control sample of the International Lung Cancer Consortium (ILCCO), comprising 27,065 individuals, we found meaningful thresholds for Iam and hiQ suitable to classify markers of poor accuracy. We demonstrate how Manhattan-like plots and moving averages of Iam and hiQ can be useful to identify regions enriched with less accurate imputed markers, whereas these regions would by missed when applying the accuracy measure info (implemented in IMPUTE2). Conclusion We recommend using Iam hiQ additional to other accuracy scores for variant filtering before stepping into the analysis of imputed GWAS data.
Genome-wide association study of more than 40,000 bipolar disorder cases provides new insights into the underlying biology
Bipolar disorder is a heritable mental illness with complex etiology. We performed a genome-wide association study of 41,917 bipolar disorder cases and 371,549 controls of European ancestry, which identified 64 associated genomic loci. Bipolar disorder risk alleles were enriched in genes in synaptic signaling pathways and brain-expressed genes, particularly those with high specificity of expression in neurons of the prefrontal cortex and hippocampus. Significant signal enrichment was found in genes encoding targets of antipsychotics, calcium channel blockers, antiepileptics and anesthetics. Integrating expression quantitative trait locus data implicated 15 genes robustly linked to bipolar disorder via gene expression, encoding druggable targets such as HTR6, MCHR1, DCLK3 and FURIN. Analyses of bipolar disorder subtypes indicated high but imperfect genetic correlation between bipolar disorder type I and II and identified additional associated loci. Together, these results advance our understanding of the biological etiology of bipolar disorder, identify novel therapeutic leads and prioritize genes for functional follow-up studies. Genome-wide association analyses of 41,917 bipolar disorder cases and 371,549 controls of European ancestry provide new insights into the etiology of this disorder and identify novel therapeutic leads and potential opportunities for drug repurposing.
Whole-genome sequencing of 234 bulls facilitates mapping of monogenic and complex traits in cattle
The 1000 bull genomes project supports the goal of accelerating the rates of genetic gain in domestic cattle while at the same time considering animal health and welfare by providing the annotated sequence variants and genotypes of key ancestor bulls. In the first phase of the 1000 bull genomes project, we sequenced the whole genomes of 234 cattle to an average of 8.3-fold coverage. This sequencing includes data for 129 individuals from the global Holstein-Friesian population, 43 individuals from the Fleckvieh breed and 15 individuals from the Jersey breed. We identified a total of 28.3 million variants, with an average of 1.44 heterozygous sites per kilobase for each individual. We demonstrate the use of this database in identifying a recessive mutation underlying embryonic death and a dominant mutation underlying lethal chrondrodysplasia. We also performed genome-wide association studies for milk production and curly coat, using imputed sequence variants, and identified variants associated with these traits in cattle.
Investigating the genetic architecture of dementia with Lewy bodies: a two-stage genome-wide association study
Dementia with Lewy bodies is the second most common form of dementia in elderly people but has been overshadowed in the research field, partly because of similarities between dementia with Lewy bodies, Parkinson's disease, and Alzheimer's disease. So far, to our knowledge, no large-scale genetic study of dementia with Lewy bodies has been done. To better understand the genetic basis of dementia with Lewy bodies, we have done a genome-wide association study with the aim of identifying genetic risk factors for this disorder. In this two-stage genome-wide association study, we collected samples from white participants of European ancestry who had been diagnosed with dementia with Lewy bodies according to established clinical or pathological criteria. In the discovery stage (with the case cohort recruited from 22 centres in ten countries and the controls derived from two publicly available database of Genotypes and Phenotypes studies [phs000404.v1.p1 and phs000982.v1.p1] in the USA), we performed genotyping and exploited the recently established Haplotype Reference Consortium panel as the basis for imputation. Pathological samples were ascertained following autopsy in each individual brain bank, whereas clinical samples were collected after participant examination. There was no specific timeframe for collection of samples. We did association analyses in all participants with dementia with Lewy bodies, and also only in participants with pathological diagnosis. In the replication stage, we performed genotyping of significant and suggestive results from the discovery stage. Lastly, we did a meta-analysis of both stages under a fixed-effects model and used logistic regression to test for association in each stage. This study included 1743 patients with dementia with Lewy bodies (1324 with pathological diagnosis) and 4454 controls (1216 patients with dementia with Lewy bodies vs 3791 controls in the discovery stage; 527 vs 663 in the replication stage). Results confirm previously reported associations: APOE (rs429358; odds ratio [OR] 2·40, 95% CI 2·14–2·70; p=1·05 × 10−48), SNCA (rs7681440; OR 0·73, 0·66–0·81; p=6·39 × 10−10), an GBA (rs35749011; OR 2·55, 1·88–3·46; p=1·78 × 10−9). They also provide some evidence for a novel candidate locus, namely CNTN1 (rs7314908; OR 1·51, 1·27–1·79; p=2·32 × 10−6); further replication will be important. Additionally, we estimate the heritable component of dementia with Lewy bodies to be about 36%. Despite the small sample size for a genome-wide association study, and acknowledging the potential biases from ascertaining samples from multiple locations, we present the most comprehensive and well powered genetic study in dementia with Lewy bodies so far. These data show that common genetic variability has a role in the disease. The Alzheimer's Society and the Lewy Body Society.
polyRAD: Genotype Calling with Uncertainty from Sequencing Data in Polyploids and Diploids
Low or uneven read depth is a common limitation of genotyping-by-sequencing (GBS) and restriction site-associated DNA sequencing (RAD-seq), resulting in high missing data rates, heterozygotes miscalled as homozygotes, and uncertainty of allele copy number in heterozygous polyploids. Bayesian genotype calling can mitigate these issues, but previously has only been implemented in software that requires a reference genome or uses priors that may be inappropriate for the population. Here we present several novel Bayesian algorithms that estimate genotype posterior probabilities, all of which are implemented in a new R package, polyRAD. Appropriate priors can be specified for mapping populations, populations in Hardy-Weinberg equilibrium, or structured populations, and in each case can be informed by genotypes at linked markers. The polyRAD software imports read depth from several existing pipelines, and outputs continuous or discrete numerical genotypes suitable for analyses such as genome-wide association and genomic prediction.
Genome-Wide Association Study and Cost-Efficient Genomic Predictions for Growth and Fillet Yield in Nile Tilapia (Oreochromis niloticus)
Fillet yield (FY) and harvest weight (HW) are economically important traits in Nile tilapia production. Genetic improvement of these traits, especially for FY, are lacking, due to the absence of efficient methods to measure the traits without sacrificing fish and the use of information from relatives to selection. However, genomic information could be used by genomic selection to improve traits that are difficult to measure directly in selection candidates, as in the case of FY. The objectives of this study were: (i) to perform genome-wide association studies (GWAS) to dissect the genetic architecture of FY and HW, (ii) to evaluate the accuracy of genotype imputation and (iii) to assess the accuracy of genomic selection using true and imputed low-density (LD) single nucleotide polymorphism (SNP) panels to determine a cost-effective strategy for practical implementation of genomic information in tilapia breeding programs. The data set consisted of 5,866 phenotyped animals and 1,238 genotyped animals (108 parents and 1,130 offspring) using a 50K SNP panel. The GWAS were performed using all genotyped and phenotyped animals. The genotyped imputation was performed from LD panels (LD0.5K, LD1K and LD3K) to high-density panel (HD), using information from parents and 20% of offspring in the reference set and the remaining 80% in the validation set. In addition, we tested the accuracy of genomic selection using true and imputed genotypes comparing the accuracy obtained from pedigree-based best linear unbiased prediction (PBLUP) and genomic predictions. The results from GWAS supports evidence of the polygenic nature of FY and HW. The accuracy of imputation ranged from 0.90 to 0.98 for LD0.5K and LD3K, respectively. The accuracy of genomic prediction outperformed the estimated breeding value from PBLUP. The use of imputation for genomic selection resulted in an increased relative accuracy independent of the trait and LD panel analyzed. The present results suggest that genotype imputation could be a cost-effective strategy for genomic selection in Nile tilapia breeding programs.
Comparing low-pass sequencing and genotyping for trait mapping in pharmacogenetics
Background Low pass sequencing has been proposed as a cost-effective alternative to genotyping arrays to identify genetic variants that influence multifactorial traits in humans. For common diseases this typically has required both large sample sizes and comprehensive variant discovery. Genotyping arrays are also routinely used to perform pharmacogenetic (PGx) experiments where sample sizes are likely to be significantly smaller, but clinically relevant effect sizes likely to be larger. Results To assess how low pass sequencing would compare to array based genotyping for PGx we compared a low-pass assay (in which 1x coverage or less of a target genome is sequenced) along with software for genotype imputation to standard approaches. We sequenced 79 individuals to 1x genome coverage and genotyped the same samples on the Affymetrix Axiom Biobank Precision Medicine Research Array (PMRA). We then down-sampled the sequencing data to 0.8x, 0.6x, and 0.4x coverage, and performed imputation. Both the genotype data and the sequencing data were further used to impute human leukocyte antigen (HLA) genotypes for all samples. We compared the sequencing data and the genotyping array data in terms of four metrics: overall concordance, concordance at single nucleotide polymorphisms in pharmacogenetics-related genes, concordance in imputed HLA genotypes, and imputation r 2 . Overall concordance between the two assays ranged from 98.2% (for 0.4x coverage sequencing) to 99.2% (for 1x coverage sequencing), with qualitatively similar numbers for the subsets of variants most important in pharmacogenetics. At common single nucleotide polymorphisms (SNPs), the mean imputation r 2 from the genotyping array was 0.90, which was comparable to the imputation r 2 from 0.4x coverage sequencing, while the mean imputation r 2 from 1x sequencing data was 0.96. Conclusions These results indicate that low-pass sequencing to a depth above 0.4x coverage attains higher power for association studies when compared to the PMRA and should be considered as a competitive alternative to genotyping arrays for trait mapping in pharmacogenetics.
GWAS of Reproductive Traits in Large White Pigs on Chip and Imputed Whole-Genome Sequencing Data
Total number born (TNB), number of stillborn (NSB), and gestation length (GL) are economically important traits in pig production, and disentangling the molecular mechanisms associated with traits can provide valuable insights into their genetic structure. Genotype imputation can be used as a practical tool to improve the marker density of single-nucleotide polymorphism (SNP) chips based on sequence data, thereby dramatically improving the power of genome-wide association studies (GWAS). In this study, we applied Beagle software to impute the 50 K chip data to the whole-genome sequencing (WGS) data with average imputation accuracy (R2) of 0.876. The target pigs, 2655 Large White pigs introduced from Canadian and French lines, were genotyped by a GeneSeek Porcine 50K chip. The 30 Large White reference pigs were the key ancestral individuals sequenced by whole-genome resequencing. To avoid population stratification, we identified genetic variants associated with reproductive traits by performing within-population GWAS and cross-population meta-analyses with data before and after imputation. Finally, several genes were detected and regarded as potential candidate genes for each of the traits: for the TNB trait: NOTCH2, KLF3, PLXDC2, NDUFV1, TLR10, CDC14A, EPC2, ORC4, ACVR2A, and GSC; for the NSB trait: NUB1, TGFBR3, ZDHHC14, FGF14, BAIAP2L1, EVI5, TAF1B, and BCAR3; for the GL trait: PPP2R2B, AMBP, MALRD1, HOXA11, and BICC1. In conclusion, expanding the size of the reference population and finding an optimal imputation strategy to ensure that more loci are obtained for GWAS under high imputation accuracy will contribute to the identification of causal mutations in pig breeding.
Optimizing Low-Cost Genotyping and Imputation Strategies for Genomic Selection in Atlantic Salmon
Genomic selection enables cumulative genetic gains in key production traits such as disease resistance, playing an important role in the economic and environmental sustainability of aquaculture production. However, it requires genome-wide genetic marker data on large populations, which can be prohibitively expensive. Genotype imputation is a cost-effective method for obtaining high-density genotypes, but its value in aquaculture breeding programs which are characterized by large full-sibling families has yet to be fully assessed. The aim of this study was to optimize the use of low-density genotypes and evaluate genotype imputation strategies for cost-effective genomic prediction. Phenotypes and genotypes (78,362 SNPs) were obtained for 610 individuals from a Scottish Atlantic salmon breeding program population (Landcatch, UK) challenged with sea lice, Lepeophtheirus salmonis. The genomic prediction accuracy of genomic selection was calculated using GBLUP approaches and compared across SNP panels of varying densities and composition, with and without imputation. Imputation was tested when parents were genotyped for the optimal SNP panel, and offspring were genotyped for a range of lower density imputation panels. Reducing SNP density had little impact on prediction accuracy until 5,000 SNPs, below which the accuracy dropped. Imputation accuracy increased with increasing imputation panel density. Genomic prediction accuracy when offspring were genotyped for just 200 SNPs, and parents for 5,000 SNPs, was 0.53. This accuracy was similar to the full high density and optimal density dataset, and markedly higher than using 200 SNPs without imputation. These results suggest that imputation from very low to medium density can be a cost-effective tool for genomic selection in Atlantic salmon breeding programs.