2013年10月2日星期三

two tools - for detecting the genetic basis of adaptation

1. DISENTANGLING THE EFFECTS OF GEOGRAPHIC AND ECOLOGICAL ISOLATION ON GENETIC DIFFERENTIATION

http://onlinelibrary.wiley.com/doi/10.1111/evo.12193/full

Populations can be genetically isolated both by geographic distance and by differences in their ecology or environment that decrease the rate of successful migration. Empirical studies often seek to investigate the relationship between genetic differentiation and some ecological variable(s) while accounting for geographic distance, but common approaches to this problem (such as the partial Mantel test) have a number of drawbacks. In this article, we present a Bayesian method that enables users to quantify the relative contributions of geographic distance and ecological distance to genetic differentiation between sampled populations or individuals. We model the allele frequencies in a set of populations at a set of unlinked loci as spatially correlated Gaussian processes, in which the covariance structure is a decreasing function of both geographic and ecological distance. Parameters of the model are estimated using a Markov chain Monte Carlo algorithm. We call this method Bayesian Estimation of Differentiation in Alleles by Spatial Structure and Local Ecology (BEDASSLE), and have implemented it in a user-friendly format in the statistical platform R. We demonstrate its utility with a simulation study and empirical applications to human and teosinte data sets.

http://genescape.ucdavis.edu/scripts-and-code/

2. INTEGRATING LANDSCAPE GENOMICS AND SPATIALLY EXPLICIT APPROACHES TO DETECT LOCI UNDER SELECTION IN CLINAL POPULATIONS

http://onlinelibrary.wiley.com/doi/10.1111/evo.12237/abstract

Uncovering the genetic basis of adaptation hinges on the ability to detect loci under selection. However, population genomics outlier approaches to detect selected loci may be inappropriate for clinal populations or those with unclear population structure because they require that individuals be clustered into populations. An alternate approach, landscape genomics, uses individual-based approaches to detect loci under selection and reveal potential environmental drivers of selection. We tested four landscape genomics methods on a simulated clinal population to determine their effectiveness at identifying a locus under varying selection strengths along an environmental gradient. We found all methods produced very low type I error rates across all selection strengths, but elevated type II error rates under “weak” selection. We then applied these methods to an AFLP genome scan of an alpine plant, Campanula barbata, and identified five highly supported candidate loci associated with precipitation variables. These loci also showed spatial autocorrelation and cline patterns indicative of selection along a precipitation gradient. Our results suggest that landscape genomics in combination with other spatial analyses provides a powerful approach for identifying loci potentially under selection and explaining spatially complex interactions between species and their environment.


2013年10月1日星期二

Computational analysis and characterization of UCE-like elements (ULEs) in plant genomes

Ultraconserved elements (UCEs), stretches of DNA that are identical between distantly related species, are enigmatic genomic features whose function is not well understood. First identified and characterized in mammals, UCEs have been proposed to play important roles in gene regulation, RNA processing, and maintaining genome integrity. However, because all of these functions can tolerate some sequence variation, their ultraconserved and ultraselected nature is not explained. We investigated whether there are highly conserved DNA elements without genic function in distantly related plant genomes. We compared the genomes of Arabidopsis thaliana and Vitis vinifera; species that diverged ∼115 million years ago (Mya). We identified 36 highly conserved elements with at least 85% similarity that are longer than 55 bp. Interestingly, these elements exhibit properties similar to mammalian UCEs, such that we named them UCE-like elements (ULEs). ULEs are located in intergenic or intronic regions and are depleted from segmental duplications. Like UCEs, ULEs are under strong purifying selection, suggesting a functional role for these elements. As their mammalian counterparts, ULEs show a sharp drop of A+T content at their borders and are enriched close to genes encoding transcription factors and genes involved in development, the latter showing preferential expression in undifferentiated tissues. By comparing the genomes of Brachypodium distachyon and Oryza sativa, species that diverged ∼50 Mya, we identified a different set of ULEs with similar properties in monocots. The identification of ULEs in plant genomes offers new opportunities to study their possible roles in genome function, integrity, and regulation.

http://genome.cshlp.org/content/22/12/2455.long




2013年9月22日星期日

Population genomics from pool sequencing

Keywords:

  • Pool sequencing;
  • High throughput sequencing;
  • Neutrality tests;
  • Composite likelihood estimators;
  • Genetic differentiation

Abstract

Next generation sequencing of pooled samples is an effective approach for studies of variability and differentiation in populations. In this paper we provide a comprehensive set of estimators of the most common statistics in population genetics based on the frequency spectrum, namely the Watterson estimator θW, nucleotide pairwise diversity II, Tajima's D, Fu and Li's D and F, Fay and Wu's H, McDonald-Kreitman and HKA tests and Fst, corrected for sequencing errors and ascertainment bias. In a simulation study, we show that pool and individual θ estimates are highly correlated and discuss how the performance of the statistics vary with read depth and sample size in different evolutionary scenarios. As an application, we reanalyze sequences from Drosophila mauritiana and from an evolution experiment in Drosophila melanogaster. These methods are useful for population genetic projects with limited budget, study of communities of individuals that are hard to isolate, or autopolyploid species.

2013年9月16日星期一

mdesci

1. http://www.medsci.cn/

2. 2013自然科学基金查询与分析系统(基础查询版)
http://www.medsci.cn/sci/nsfc.do

3. MedSci 2013年期刊智能查询系统(2012年度)
http://www.medsci.cn/sci/submit.asp

4. 论文服务
http://www.medsci.cn/list.asp?classid=110

public library of bioinformatics

1. http://www.plob.org/
public library of bioinformatics

2. http://www.bioask.net/

2013年9月14日星期六

forest plot

https://mcfromnz.wordpress.com/2012/11/06/forest-plots-in-r-ggplot-with-side-table/#more-356


BroadE Workshop 2013 July 9-10

http://www.broadinstitute.org/gatk/guide/events?id=3093#materials

This workshop covered the core steps involved in calling variants with the Broad’s Genome Analysis Toolkit, using the “Best Practices” developed by the GATK team. View the workshop materials to learn why each step is essential to the calling process, what are the key operations performed on the data at each step, and how to use the GATK tools to get the most accurate and reliable results out of your dataset.

Workshop materials


 - Day 1 - Opening remarks

 -  - Introduction to Next Generation Sequence Analysis

 -  - Introduction to the GATK

 -  - Mapping and duplicate marking (data pre-processing)

 -  - Local realignment around indels
RTC IR

 -  - Base quality score recalibration (BQSR)
BR PR

 -  - Compression with ReduceReads
RR

 - Day 2 - Opening remarks

 -  - Variant calling
UG HC

 -  - Variant quality score recalibration (VQSR)
VR AR

 -  - Genotype phasing and refinement
PBT RBP

 -  - Functional annotation
VA

 -  - Analyzing variant calls
SV CV VE

 - Introduction to Parallelism (video not available yet)
NT NCT Q



Supplemental materials


 -  - GenomeSTRiP: Discovery and genotyping of deletions

 - XHMM: Discovery and genotyping of copy number variation from exome read depth (PDF not available for download yet)

2013年9月9日星期一

2013年龙星计划之生物信息学

http://yixf.name/2013/09/04/%E8%8D%902013%E5%B9%B4%E9%BE%99%E6%98%9F%E8%AE%A1%E5%88%92%E4%B9%8B%E7%94%9F%E7%89%A9%E4%BF%A1%E6%81%AF%E5%AD%A6/

课程主页

课件下载

课程视频

课程简介

  • Day 1. Background. Basic Statistics. Introduce deep sequencing data. Motivational examples.
  • Day 2. Analyze RNA-seq data and small RNA-seq data.
  • Day 3. DNA methylation, Integration with other data types.
  • Day 4. Analyze ChIP-seq data on transcription factors and histone modifications. Integration with other sequencing data types.
  • Day 5. Analyze DNase-seq data and MNase-seq data. Integration with other data types.

实验内容

  • Day 1. Background. Basic Statistics. Introduce deep sequencing data. Motivational examples.
  • Day 2. Analyze ChIP-seq data on transcription factors and histone modifications. Integration with other sequencing data types.
  • Day 3. Analyze RNA-seq data and small RNA-seq data
  • Day 4. Analyze DNase-seq data and MNase-seq data. Integration with other data types
  • Day 5. DNA methylation, Integration with other data types

MOSAIK: A hash-based algorithm for accurate next-generation sequencing read mapping

http://arxiv.org/pdf/1309.1149v1.pdf

MOSAIK is a stable, sensitive and open-source program for mapping second and third-generation
sequencing reads to a reference genome. Uniquely among current mapping tools, MOSAIK can align
reads generated by all the major sequencing technologies, including Illumina, Applied Biosystems SOLiD,
Roche 454, Ion Torrent and Pacific BioSciences SMRT. Indeed, MOSAIK was the only aligner to
provide consistent mappings for all the generated data (sequencing technologies, low-coverage and
exome) in the 1000 Genomes Project. To provide highly accurate alignments, MOSAIK employs a hash
clustering strategy coupled with the Smith-Waterman algorithm. This method is well-suited to capture
mismatches as well as short insertions and deletions. To support the growing interest in larger structural
variant (SV) discovery, MOSAIK provides explicit support for handling known-sequence SVs, e.g.
mobile element insertions (MEIs) as well as generating outputs tailored to aid in SV discovery. All
variant discovery benefits from an accurate description of the read placement confidence. To this end,
MOSAIK uses a neural-net based training scheme to provide well-calibrated mapping quality scores,
demonstrated by a correlation coefficient between MOSAIK assigned and actual mapping qualities
greater than 0.98. In order to ensure that studies of any genome are supported, a training pipeline is
provided to ensure optimal mapping quality scores for the genome under investigation. MOSAIK is
multi-threaded, open source, and incorporated into our command and pipeline launcher system GKNO
(http://gkno.me).

2013年8月23日星期五

一个新生境下物种形成的实例

Rapid speciation with gene flow following the formation of Mount Etna

http://gbe.oxfordjournals.org/content/early/2013/08/23/gbe.evt127.full.pdf+html

Environmental or geological changes can create new niches which drive ecological species divergence without the immediate cessation of gene flow. However, few such cases have been characterised. On the recently formed volcano, Mt. Etna, Senecio aethnensis and S. chrysanthemifolius inhabit contrasting environments of high and low altitude respectively. They have very distinct phenotypes, despite hybridising promiscuously, and thus may represent an important example of ecological speciation ‘in action’, possibly as a response to the rapid geological changes which Mt. Etna has recently undergone. To elucidate the species' evolutionary history, and help establish the species as study system for speciation genomics, we sequenced the transcriptomes of the two Etnean species, and the outgroup, S. vernalis, using Illumina sequencing. Despite the species' substantial phenotypic divergence, synonymous divergence between the high- and low-altitude species was low (dS = 0.016 ± 0.017 [SD]). A comparison of species divergence models with and without gene flow provided unequivocal support in favor of the former and demonstrated a recent time of species divergence (153,080 ya ± 11,470[SE]) that coincides with the growth of Mount Etna to the altitudes which separate the species today. Analysis of dN/dSrevealed wide variation in selective constraint between genes, and evidence that highly expressed genes, more ‘multifunctional’ genes and those with more paralogues were under elevated purifying selection. Taken together, these results are consistent with a model of ecological speciation, potentially as a response to the emergence of a new, high altitude niche as the volcano grew.

2013年8月18日星期日

Alternative forms for genomic clines

http://onlinelibrary.wiley.com/doi/10.1002/ece3.609/full

Understanding factors regulating hybrid fitness and gene exchange is a major research challenge for evolutionary biology. Genomic cline analysis has been used to evaluate alternative patterns of introgression, but only two models have been used widely and the approach has generally lacked a hypothesis testing framework for distinguishing effects of selection and drift. I propose two alternative cline models, implement multivariate outlier detection to identify markers associated with hybrid fitness, and simulate hybrid zone dynamics to evaluate the signatures of different modes of selection. Analysis of simulated data shows that previous approaches are prone to false positives (multinomial regression) or relatively insensitive to outlier loci affected by selection (Barton's concordance). The new, theory-based logit-logistic cline model is generally best at detecting loci affecting hybrid fitness. Although some generalizations can be made about different modes of selection, there is no one-to-one correspondence between pattern and process. These new methods will enhance our ability to extract important information about the genetics of reproductive isolation and hybrid fitness. However, much remains to be done to relate statistical patterns to particular evolutionary processes. The methods described here are implemented in a freely available package “HIest” for the R statistical software (CRAN; http://cran.r-project.org/).

Theoretical Evolutionary Genetics - draft text

1. http://evolution.genetics.washington.edu/pgbook/pgbook.html

This would be a very good book on population genetics.

2. Evolution and Selection of Quantitative Traits by Bruce Walsh and Michael Lynch. While this book is in draft form it is available from Bruce Walsh's web page at: http://nitro.biosci.arizona.edu/zbook/NewVolume_2/newvol2.html (Bruce Walsh's web page is in general a fantastic source of information on all things population/quantitative genetics).

3. from Withlock in UBC

2013年8月12日星期一

perl tutorial

http://web.guru99.com/perl-tutorials/

hybrid zone

hybrid zones allow us:
(1) to quantify the genetic differences responsible for speciation,
(2) to measure the diffusion of genes between diverging taxa,
(3) to understand the spread of alternative adaptations.

The genomic impacts of drift and selection for hybrid performance

1. http://arxiv.org/abs/1307.7313

Modern maize breeding relies upon selection in inbreeding populations to improve the performance of cross-population hybrids. The United States Department of Agriculture - Agricultural Research Service reciprocal recurrent selection experiment between the Iowa Stiff Stalk Synthetic (BSSS) and the Iowa Corn Borer Synthetic No. 1 (BSCB1) populations represents one of the longest standing models of selection for hybrid performance. To investigate the genomic impact of this selection program, we used the Illumina MaizeSNP50 high-density SNP array to determine genotypes of progenitor lines and over 600 individuals across multiple cycles of selection. Consistent with previous research (Messmer et al., 1991; Labate et al., 1997; Hagdorn et al., 2003; Hinze et al., 2005), we found that genetic diversity within each population steadily decreases, with a corresponding increase in population structure. High marker density also enabled the first view of haplotype ancestry, fixation and recombination within this historic maize experiment. Extensive regions of haplotype fixation within each population are visible in the pericentromeric regions, where large blocks trace back to single founder inbreds. Simulation attributes most of the observed reduction in genetic diversity to genetic drift. Signatures of selection were difficult to observe in the background of this strong genetic drift, but heterozygosity in each population has fallen more than expected. Regions of haplotype fixation represent the most likely targets of selection, but as observed in other germplasm selected for hybrid performance (Feng et al., 2006), there is no overlap between the most likely targets of selection in the two populations. We discuss how this pattern is likely to occur during selection for hybrid performance, and how it poses challenges for dissecting the impacts of modern breeding and selection on the maize genome.

How do I match orthologues in one species to another, genome scale

http://www.biostars.org/p/569/

how to detect ortholog among species.

Subset of heat-shock transcription factors required for the early response of Arabidopsis to excess light

1. http://www.sciencedaily.com/releases/2013/08/130806132939.htm

2. http://www.pnas.org/content/early/2013/07/31/1311632110

How Increasing CO2 and Temperatures Affect Plant Development

1. http://www.sciencedaily.com/releases/2013/07/130731225931.htm

2. http://www.nature.com/ncomms/2013/130731/ncomms3145/full/ncomms3145.html

Elevated levels of CO2 and temperature can both affect plant growth and development, but the signalling pathways regulating these processes are still obscure. MicroRNAs function to silence gene expression, and environmental stresses can alter their expressions. Here we identify, using the small RNA-sequencing method, microRNAs that change significantly in expression by either doubling the atmospheric CO2 concentration or by increasing temperature 3–6 °C. Notably, nearly all CO2-influenced microRNAs are affected inversely by elevated temperature. Using the RNA-sequencing method, we determine strongly correlated expression changes between miR156/157 and miR172, and their target transcription factors under elevated CO2 concentration. Similar correlations are also found for microRNAs acting in auxin-signalling, stress responses and potential cell wall carbohydrate synthesis. Our results demonstrate that both CO2 and temperature alter microRNA expression to affect Arabidopsis growth and development, and miR156/157- and miR172-regulated transcriptional network might underlie the onset of early flowering induced by increasing CO2.