2013年2月12日星期二

PROVEAN, SIFT - variant annotation

PROVEAN (Protein Variation Effect Analyzer), which provides a generalized approach to predict the functional effects of protein sequence variations including single or multiple amino acid substitutions, and in-frame insertions and deletions. The PROVEAN tool is available online at http://provean.jcvi.org

SIFT predicts whether an amino acid substitution affects protein function. SIFT prediction is based on the degree of conservation of amino acid residues in sequence alignments derived from closely related sequences, collected through PSI-BLAST. SIFT can be applied to naturally occurring nonsynonymous polymorphisms or laboratory-induced missense mutations.

FastQC - on SAM/BAM and FastQ


FastQC aims to provide a simple way to do some quality control checks on raw sequence data coming from high throughput sequencing pipelines. It provides a modular set of analyses which you can use to give a quick impression of whether your data has any problems of which you should be aware before doing any further analysis.
The main functions of FastQC are
  • Import of data from BAM, SAM or FastQ files (any variant)
  • Providing a quick overview to tell you in which areas there may be problems
  • Summary graphs and tables to quickly assess your data
  • Export of results to an HTML based permanent report
  • Offline operation to allow automated generation of reports without running the interactive application

tools for variant analysis of NGS data


A survey of tools for variant analysis of next-generation genome sequencing data



http://bib.oxfordjournals.org/content/early/2013/01/21/bib.bbs086.full



#########################################################################

Analysis pipelines and workflow systems

We evaluated three analytical pipelines (‘HugeSeq’, ‘SIMPLEX’ and ‘TREAT’) and three workflow systems (‘Galaxy’, ‘LONI’ and ‘Taverna’).
‘HugeSeq’ [139] is a fully integrated pipeline for NGS analysis from aligning reads to the identification and annotation of variants (SNPs and INDELs for whole-genome and whole-exome sequencing data as well as CNVs and SVs for whole-genome data only). It consists of three main parts: (i) preparing and aligning reads, (ii) combining and sorting reads for parallel processing of variant calling and (iii) variant calling and annotating. The pipeline accepts as an input reads in FASTA or FASTQ format and outputs identified variants in VCF format, and SVs and CNVs in GFF format. Identified variants are further processed with ANNOVAR to include additional annotations.



file conversion perl scripts

1.convert gff files to gtf files
convert_chado_gff_to_gtf.pl

2. convert emble files to fasta files
embl_to_fasta.pl

Inferring transcript sequences from a gff file

http://avrilomics.blogspot.ca/2013/01/inferring-transcript-sequences-from-gff.html

get_spliced_transcripts_from_gff.pl that does this, based on an input gff file of gene predictions, and an input fasta file of chromosome/scaffold sequences.

2013年2月11日星期一

Script to parse fasta headers

http://www.biostars.org/p/62884/


Assuming that you do not have spaces in your sequences you can try that:
sed -e '/^>/ s/ .*//' mybigfile
Awk should work too (even if your sequences contain spaces):
awk '{print /^>/ ? $1 : $0}' mybigfile
and just for fun, a pure bash version:
while read l ; do echo "${l%% *}" ; done < mybigfile

rOpenSci is a collaborative effort to develop R-based tools for facilitating Open Science

At rOpenSci we are creating packages that allow access to data repositories through the R statistical programming environment that is already a familiar part of the workflow of many scientists. We hope that our tools will not only facilitate drawing data into an environment where it can readily be manipulated, but also one in which those analyses and methods can be easily shared, replicated, and extended by other researchers. While all the pieces for connecting researchers with these data sources exist as disparate entities, our efforts will provide a unified framework that will be quickly connect researchers to open data.

http://ropensci.org/

加莱义民(The Burghers of Calais )

http://baike.baidu.com/view/295049.htm


rentrez - search or download data from various NCBI databases.


rentrez provides functions that work with the NCBI eutils to search or download data from various NCBI databases.
The package hasn't been thoroughly tested yet, but the functions for each of the Eutils functions are implimented. If you try the package and find bugs please let me know.

2013年2月10日星期日

genetic diversity and the rate of recombination


Genome-wide analysis in chicken reveals that local levels of genetic diversity are mainly governed by the rate of recombination


Background

Polymorphism is key to the evolutionary potential of populations. Understanding which factors shape levels of genetic diversity within genomes forms a central question in evolutionary genomics and is of importance for the possibility to infer episodes of adaptive evolution from signs of reduced diversity. There is an on-going debate on the relative role of mutation and selection in governing diversity levels. This question is also related to the role of recombination because recombination is expected to indirectly affect polymorphism via the efficacy of selection. Moreover, recombination might itself be mutagenic and thereby assert a direct effect on diversity levels.

Results

We used whole-genome re-sequencing data from domestic chicken (broiler and layer breeds) and its wild ancestor (the red jungle fowl) to study the relationship between genetic diversity and several genomic parameters. We found that recombination rate had the largest effect on local levels of nucleotide diversity. The fact that divergence (a proxy for mutation rate) and recombination rate were negatively correlated argues against a mutagenic role of recombination. Furthermore, divergence had limited influence on polymorphism.

Conclusions

Overall, our results are consistent with a selection model, in which regions within a short distance from loci under selection show reduced polymorphism levels. This conclusion lends further support from the observations of strong correlations between intergenic levels of diversity and diversity at synonymous as well as non-synonymous sites. Our results also demonstrate differences between the two domestic breeds and red jungle fowl, where the domestic breeds show a stronger relationship between intergenic diversity levels and diversity at synonymous and non-synonymous sites. This finding, together with overall lower diversity levels in domesticates compared to red jungle fowl, seem attributable to artificial selection during domestication.

homologous and non-homologous recombination in genome evolution


Impact of homologous and non-homologous recombination in the genomic evolution of Escherichia coli


Genomicus: five genome browsers for comparative genomics in eukaryota


Genomicus: five genome browsers for comparative genomics in eukaryota


Genomicus (http://www.dyogen.ens.fr/genomicus/) is a database and an online tool that allows easy comparative genomic visualization in >150 eukaryote genomes. It provides a way to explore spatial information related to gene organization within and between genomes and temporal relationships related to gene and genome evolution. For the specific vertebrate phylum, it also provides access to ancestral gene order reconstructions and conserved non-coding elements information. We extended the Genomicus database originally dedicated to vertebrate to four new clades, including plants, non-vertebrate metazoa, protists and fungi. This visualization tool allows evolutionary phylogenomics analysis and exploration. Here, we describe the graphical modules of Genomicus and show how it is capable of revealing differential gene loss and gain, segmental or genome duplications and study the evolution of a locus through homology relationships.


Evolutionary Rate and Duplicability in the Arabidopsis thaliana Protein–Protein Interaction Network

http://gbe.oxfordjournals.org/content/4/12/1263.full

Genes show a bewildering variation in their patterns of molecular evolution, as a result of the action of different levels and types of selective forces. The factors underlying this variation are, however, still poorly understood. In the last decade, the position of proteins in the protein–protein interaction network has been put forward as a determinant factor of the evolutionary rate and duplicability of their encoding genes. This conclusion, however, has been based on the analysis of the limited number of microbes and animals for which interactome-level data are available (essentially, Escherichia coli, yeast, worm, fly, and humans). Here, we study, for the first time, the relationship between the position of proteins in the high-density interactome of a plant (Arabidopsis thaliana) and the patterns of molecular evolution of their encoding genes. We found that genes whose encoded products act at the center of the network are more evolutionarily constrained than those acting at the network periphery. This trend remains significant when potential confounding factors (gene expression level and breadth, duplicability, function, and length of the encoded products) are controlled for. Even though the correlation between centrality measures and rates of evolution is generally weak, for some functional categories, it is comparable in strength to (or even stronger than) the correlation between evolutionary rates and expression levels or breadths. In addition, genes encoding interacting proteins in the network evolve at relatively similar rates. Finally, Arabidopsis proteins encoded by duplicated genes are more highly connected than those encoded by singleton genes. This observation is in agreement with the patterns observed in humans, but in contrast with those observed in E. coli, yeast, worm, and fly (whose duplicated genes tend to act at the periphery of the network), implying that the relationship between duplicability and centrality inverted at least twice during eukaryote evolution. Taken together, these results indicate that the structure of the A. thaliana network constrains the evolution of its components at multiple levels.

2013年2月8日星期五

QIIME scripts - tools for quantitative Ecology

http://qiime.wordpress.com/

http://qiime.org/scripts/

Sequence assembly with MIRA3, a Swiss knife like tool


Sequence assembly with MIRA3



MIRA is a whole genome shotgun and EST sequence assembler for Sanger, 454, Solexa (Illumina) and PacBio data. It can be seen as a Swiss army knife of sequence assembly developed and used in the past 12 years to get assembly jobs done efficiently - and especially accurately. That is, without actually putting too much manual work into finishing the assembly.

References: 
  • Chevreux et al. (1999) Genome Sequence Assembly Using Trace Signals and Additional Sequence Information Computer Science and Biology: Proceedings of the German Conference on Bioinformatics (GCB) 99, pp. 45-56.
  • Chevreux et al. (2004) Using the miraEST Assembler for Reliable and Automated mRNA Transcript Assembly and SNP Detection in Sequenced ESTs Genome Research 2004. 14:1147-1159.

2013年2月7日星期四

perl scripts on systematic and evolutionary study

http://www.molekularesystematik.uni-oldenburg.de/33997.html#Sequences


SupertreesSequencesEvoDevoMiscellaneous

SUPERTREES (AND TREES IN GENERAL)

v.1.2.2
Determines the bootstrap frequencies for a given phylogenetic tree based on the results of a bootstrap analysis.
v.1.3.3
Adds branch length information to a NEXUS-formatted tree description corresponding to divergence times estimates of the nodes to create an ultrametric "chronograph". Will interpolate any missing dates using the log-log formula of Purvis (1995; see Bininda-Emonds et al., 1999 for the reference) and can correct for negative branch lengths.
v.1.1
Labels internal nodes of a NEXUS-formatted tree description with names of higher-level taxa according to a user-input taxonomy. These labels can be viewed in programs such as TreeView.
v.1.0
Adds node numbers to a NEXUS-formatted tree description for presentation purposes. These labels can be viewed in programs such asTreeView.
v.1.2.1
Calculates the number of unique clades / partitions for pairs of trees pruned to their common taxon set. Will optionally calculate a value weighted according to a measure of nodal support (given as a branch length), and can use this to calculate if two topologically identical trees have statistically different support values.
v.1.2.1
Calculates the qualitative support for the clades present in a supertree relative to those in the source trees contributing to it. Described inBininda-Emonds (2003) with a modified version (rQS) described inPrice et al. (2005).
Prachi Shah and Davide Pisani of Penn State University have kindly ported (an older version of) this program as a DOS executable file. It might only operate using the default parameters, however.
v.2.2
Derives relative branch length formulas for dating a supertree from one or more gene trees according to the "local molecular clock" procedure in Purvis (1995; see Bininda-Emonds et al., 1999 for the reference). Dates can be relative to either ancestral or daughter nodes.
v.1.0a
Uses PAUP* to reverse engineer the source trees from a NEXUS-formatted MRP data matrix and store in a single tree file. Requires, naturally, that the character boundaries of the source trees are specified (as nexus-format CHARSETs).
v.1.2.1
Converts a NEXUS-formatted treefile into a NEXUS-formatted data file ready for analysis. Incorporates both standard and Purvis MRP coding, and allows source trees to be coded as either rooted or unrooted (the latter as described in Bininda-Emonds et al., 2005).
v.2.1
Standardizes the taxon names in a set of source trees according to a user-input reference taxonomy and synonomy list. Mismatches are flagged for the user to correct. Note that it cannot account for branch-length or support information in the trees. Described in Bininda-Emondset al. (2004).
(Note: versions 1.0.x had serious bugs and should not be used!)
taxonoTree.plv.1.0Constructs the tree associated with a hierarchical taxonomy presented in a tab-delimited text file.
v.1.0
Converts a NEXUS-formatted treefile into a PHYLIP-formatted treefile or standalone NEXUS-formatted data file.
treePruner.plv.1.0Prunes the trees in a NEXUS-formatted treefile to their common taxon set, of specific user-input taxa, or both. Can optionally retain support values in the trees.

SupertreesSequencesEvoDevoMiscellaneous

SEQUENCES (AND DATA MINING)

autoMT.plv.1.0Allows for batch testing of the optimal model of evolution for a series of sequence files. Model testing can be performed using either ModelTEST with PAUP* or MrAIC.pl with PHYML. The applicability of the molecular clock can also be tested using the ModelTEST / PAUP* combination.
batchPHYML.plv.1.0Provides a wrapper around PHYML to easily perform sequential analyzes on a set of data matrices specfied by the user in a tab-delimited text file.
batchRAXML.plv.1.1.1Provides a wrapper around RAxML to easily analyze a set of data files according to a common set of the search criteria. Also organizes the RAxML output into a set of subdirectories. Compatible with RAxML-VI-HPC v2.2.3.
v.2.0
Mines all gene sequences from a GenBank output file according to annotations provided in each accession. As such, it is limited by the accuracy of the information given in the accession and uses a restricted library of gene synonyms. However, it can often mine more evolutionarily divergent sequences and better account for paralogs than can a BLAST-based search.
v.1.0a
Crude program to count number of sequences for a given gene in a GenBank download. Does not correct for differences in spelling, etc.
moleRat.plv1.0Calculates rates of evolution along the branches of one or more (gene) trees with respect to a dated reference tree, both for each tree individually and across the set of trees as a whole. Also identifies branches and clades that are evolving significantly differently from the overall average or that have changed their rate significantly with respect to an ancestral reference point (as determined using a paired Student's t-test and a paired Fisher's sign test). Described in Bininda-Emonds (2007).
seqCat.plv1.0Creates an interleaved nexus-formatted supermatrix of individual data matrices (in any of fasta, NEXUS, PHYLIP, or Se-Al formats).
v.1.0.2
Processes an aligned DNA sequence data set to retain only 1) those sequences with a minimum level of pairwise overlap and 2) the five most diverse (and longest) sequences for taxa with greater than five sequences. Data can be input in any of fasta, NEXUS, PHYLIP, or Se-Al formats.
seqConverter.plv.1.2Convert between some commonly used file formats (fasta, NEXUS, PHYLIP, or Se-Al) as well as performs simple data transformations (e.g., modify gaps, translate to amino acids, convert to haplotype data). Can also batch convert all programs of a specified file type in the working directory. A program that recognizes more file formats is sreformat, part of the HMMERpackage.
v.1.2
Facilitates the multiple alignment of protein-coding DNA sequences by aligning the amino acids sequences they specify. Data can be input in and output to any of fasta, NEXUS, PHYLIP, or Se-Alformats. Requires a local copy of ClustalW. Described in Bininda-Emonds (2005).

SupertreesSequencesEvoDevoMiscellaneous

EVODEVO

The Parsimov package, as described in Jeffery et al. (2005). All programs and example files were written by Jonathan Jeffery.
v.1.0.7g
Implements event-pair "parsimony cracking" as described inJeffery et al. (2005).
v.1.0.3b
Takes a Parsimv7g.pl output file and replaces the PAUP* character numbers with more readable character names according to a user-specified text list (e.g., CellType.txt, based of cell-line characters in spiralians).  For the latter, each line contains the character number (in ascending order) and character name, separated by a tab.
A lot of replacing can be automated using a batch file (e.g., isReplaceBat.txt).  The batch file contains the name of the output file to have its character numbers replaced and the name of the text-list to use, separated by a tab.
n/a
Creates a PAUP* command file to describe each tree in memory under ACCTRAN and DELTRAN optimizations (saving each as separate log files) plus a Parsimv7g.pl batch file (e.g.,ParsBatch.txt) to crack each of the PAUP* log files produced.  Handy for big jobs.
n/a
An example of a batch file -- if you have several log files to work through (e.g., ACCTRAN or DELTRAN optimizations, different topologies, etc), you can use this to get Parsimv7g.pl to run through each in turn.  The contents are, for each log-file you want to analyze: the name of a log file, the path where you want its output files written, whether to use all [a] or unambiguous [u] changes, whether to use a thorough search if feasible [y/n] and whether to clean-up the "working" files as it goes along [y/n] (useful for big data sets where temp files can reach 100s of MB).  This data must be tab-separated on a new line for each log file.
Other EvoDevo programs (written by me). Note: except for GamSim.pl, these programs are somewhat dated and run in MacPerl only.
v.1.0.2
Implements event-pair cracking as described in Jeffery et al. (2002).
BreakPoint.pl EventPair.pl 
JuncCode.pl
v.1.0
These three programs will encode developmental timing data according to three different coding schemes: breakpoint distances, event-pairing, and junction coding. Each program will output a NEXUS-formatted data file ready for analysis.
v.1.0
Infers the developmental sequence of a hypothetical ancestor on a cladogram, following the procedure described in Jeffery et al. (2002).
v.1.0a
Calculates whether the events in a given developmental sequence are evenly distributed or not. Described in Bininda-Emonds et al. (2003).

SupertreesSequencesEvoDevoMiscellaneous

MISCELLANEOUS

PerlEQ.plv.1.0b12A program, written by Jonathan Jeffery, that performs Safe Taxonomic Reduction (developed by Mark Wilkinson) to identify taxa that can be safely removed from a phylogenetic analysis because they are essentially "redundant" with other taxa. I have produced a modified version that allows all input options to be specified from the command line, including a new option to suppress output of the data matrix and character diagnostics to the html file (to keep the size of this file down somewhat). 
(Note: the program is no longer being actively maintained.)
v.1.0.9a
Creates a PAUP* batch file with the necessary instructions to perform a parsimony ratchet analysis.
reverseSTR.plv.1.0Re-includes taxa, where this is unequivocally possible, to a NEXUS-formatted tree description derived from a analysis using Safe Taxonomic Reduction. Requires the output of STRindexer.pl.
STRindexer.plv.1.1Parses the html output file of PerlEQ to identify taxa that can be safely removed from an analysis and then potentially unequivocally re-included (using reverseSTR.pl). Essentially, these are taxa conforming to the PerlEQ category C*.
STRedundancy.plv.1.0Identifies characters in a data matrix that will become redundant (i.e., duplicate others or become uninformative) upon deletion of a specified set of taxa. It is geared largely towards STR analyses of MRP matrices.
v.1.0
Converts Mac and DOS-style line breaks to Unix-style ones. This is a holdover from when my scripts required Unix-style line breaks.Here is an even better program (for Mac OS X) with drag-n-drop and the works.

on calculation of dN/dS ratio

1. HyPhy
http://www.hyphy.org/w/index.php/Main_Page

HyPhy is an open-source software package for the analysis of genetic sequences using techniques in phylogenetics, molecular evolution, and machine learning. It features a complete graphical user interface (GUI) and a rich scripting language for limitless customization of analyses. Additionally, HyPhy features support for parallel computing environments (via message passing interface) and it can be compiled as a shared library and called from other programming environments such as Python or R. HyPhy has over 7000 registered users and has been cited in over 600 peer-reviewed publications (Google Scholar). Continued development of HyPhy is currently supported in part by an NIGMS R01 award 1R01GM093939.

2. web-based tools

http://www.datamonkey.org/


3. JCoDA: a tool for detecting evolutionary selection

http://www.biomedcentral.com/1471-2105/11/284/

4. IDEA: Interactive Display for Evolutionary Analyses
http://www.biomedcentral.com/1471-2105/9/524

5. SNAP - web-based tools
http://hcv.lanl.gov/content/sequence/SNAP/SNAP.html

multiple sequence alignment - DIALIGN-TX


Introduction:

DIALIGN-TX is a substantial improvement of DIALIGN-T that combines greedy and progressive alignment strategies in a new algorithm which is now available for download. Further information can be found in:

Further Information and Citation

More detailed descriptions of the methods can be found in:
Research work using DIALIGN-TX/DIALIGN-T should cite the above mentioned publications.


http://dialign-tx.gobics.de/

http://dialign-tx.gobics.de/submission?type=dna

2013年2月6日星期三

Plant Resistance Gene -- R genes Wiki


PRGdb 2.0: towards a community-based database model for the analysis of R-genes in plants


A list of bioinformatics courses

1. http://ged.msu.edu/angus/bioinformatics-courses.html

Below is a list of bioinformatics courses that I’ve run across over time (many of them sent to me on Twitter). They should be useful places to go for training materials and/or actual training :). Please post any additional ones in the comments at the bottom and I’ll update these links. Thanks! –titus
NESCent Academy – Next-generation sequencing in evolutionary biology https://academy.nescent.org/wiki/Main_Page
Microbial Metagenomics 2012 @ Michigan State – http://metagenomics.wikidot.com/
Cold Spring Harbor / Advanced Sequencing Technologies – http://meetings.cshl.edu/courses/c-seqtech12.shtml
Indiana University (semester-long) - http://compbio.iupui.edu/group/1/pages/next_gen_course
Bioinformatics for Cancer Genomics / Toronto, CA – http://bioinformatics.ca/workshops/2012/bioinformatics-cancer-genomics-bicg
Informatics on High Throughput Sequencing Data / Toronto, CA – http://bioinformatics.ca/workshops/2012/informatics-high-throughput-sequencing-data
European Sequence Data Analysis Training School – http://www.basgen.nl/sdac/
Materials for an Australian NGS workshop – https://github.com/nathanhaigh/ngs_workshop#readme
2.