Thursday, September 29, 2016

Initial Project Planning

Network Diagram for the next 3 months

Updated Gantt Chart


Over the next three months I will be getting all the preliminary steps of creating a bioinformatics website out of the way. I will learn how to code in Javascript and CSS and then begin gathering the data that the website will organize. Finally, I will start the laying the foundations of the website.

Initial Project Planning:
Identify the Problem:
Lyme disease is constantly evolving and new strains are emerging. This makes it increasingly hard to treat because the strains are evolving to evade antibiotics. We need to sequence all existing strains and find which sections are increasing virulence and allowing the strains to evade the immune system and which sections are conserved.
Prior Art Survey:
There is a bioinformatics site which has compiled the sequences of 35 genomes of 8 LB species and 7 RF species. It is a manually curated site which includes comparisons of phylogeny, synteny and sequence alignments of orthologous sequences and intergenic spacers. (http://borreliabase.org/)
Literature Review: 
Can be found in previous posts.
Solution Space Exploration: 
Since a site already exists with with MSA, phylogeny and synteny. They do not have a complete genome for most strains. They also do not have the functions of each gene.
High Level Design: 
I need to get in contact with the people who wrote the paper I read over the summer about lime disease before I start my project.

Tools:
If I plan to work off of the existing database for lyme disease I will need just a computer and possibly access to the NCBI database, in order to read the articles on lyme disease strains. I could access these through the library.
If I plan to work with sequencing that have not been sequenced yet, then I will need access to a lab or some sort of collaboration with a lab to sequence the rest of the genome.
I will need a mentor who knows about the lyme disease genome

Budget:
Most likely 0$

http://www.novusbio.com/diseases/lyme-disease
http://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-15-233
Bioinformatics Slide:
https://docs.google.com/presentation/d/17VzdZdIEADFSacXulYF5ocCSfWaYhgboPMJYiH3iaGA/edit?usp=sharing

Thursday, September 15, 2016

Bioinformatics Week 9

Progress on project compiled:
https://docs.google.com/presentation/d/17VzdZdIEADFSacXulYF5ocCSfWaYhgboPMJYiH3iaGA/edit?usp=sharing


Saturday, September 3, 2016

Bioinformatics Week 8

Next generation sequencing: Generates large amount of sequences. 
The price of sequencing has dropped significantly since the start of sequencing. 
454 Pyrosequecning

  • DNA is broken into smaller sequences. 
  • Attach single peice of DNA is attatched to a bead
  • Single pieces of DNA on the beads are then applified
  • Sequencing reactions occur and the sequencing data is created. 
Illumina Sequencing 
  • DNA is broken into smaller sequences 
  • Clusters of the DNA are created by adding adapters(stick together non congrous DNA)
  • Clusters are amplified in PCR
  • Then you sequence each individual cluster by adding in primers and nucleotides which are tagged with different colored florescences. 
The current sequencing technologies fall short in some aspects. First, they all break the DNA into smaller segments and then sequence them. So you are sequencing small parts and it's harder to see the bigger picture. You have to try to piece back together the small sequences to figure out what the whole piece of DNA was. When there are repetitive regions within DNA it becomes more complicated to piece back the DNA sequences once you have sequenced it. You don't know where these sequences were located on the DNA. Putting the sequences back together is called assembly. There are different programs which assemble the sequences (ABySS, SOAPdenovo, Velvet). When you are given an assembly you are also given an assembly score. This assembly score comes with the number of contigs (overlapping sequences). The fewer contigs the better. 

RNA-seq: Sequences mRNA. Can identify if alternative splicing events have occurred (removing introns, not removing introns, etc)
  • DNA is broken into smaller sequences 
  • Clusters of the DNA are created by adding adapters(stick together non congruous DNA)
  • Sequence it 
Metagenomics: Studying the sequences of microbial organism in their natural environments by just taking samples of their environment (soil, water, etc) and then you sequence it using next generation sequencing. Before you would have to remove the organism from its natural environment and then culture it in an artificial lab environment which could alter your results. 


Friday, September 2, 2016

Bioinformatics Week 7

Selection analysis: Tries to identify whether natural selection is occurring or not as well as if the selection is to get rid of a certain sequence or promote a sequence. Then how this selection affects the species. 
3 types of selection

  • Negative: Removes a detrimental mutation. 
  • Positive: Bennificial mutation promoted. 
  • Balancing/Diversfying seleciton: Favors the maintenance of multiple variations of a sequence in a diverse environment. 
When measuring selection there are two standard techniques
  • Tajima's D: Based on a few standard principles. The genetic regions near a region which is being selected for get dragged into the selection process (genetic hitchhiking). The length of the genetic region is dependent on the rate of recombination. Theta equals to Pi when no selection is occurring. When theta is greater than Pi then positive selection is occurring. When theta is less than Pi then balancing selection is occurring. Advantages: Good however can be fooled by other factors (ex:bottleneck mistaken for positive selection). Therefore after you preform Tajima's D you must preform other tests to single out chances of error in your data. 
  • dN/dS: This calculates the ratio of nonsynonmous mutations (change the protein) to synonemous mutations (doesn't affect protein). Very good for figuring out if there are specific sites which are being selected for and then figure out which codons are being selected for. No selection dN/dS=1, negative selection dN/dS<1, positive selection dN/dS>1. 
There are two popular ways of quantifying selection (measuring variation)
  • Theta: Based on the number of variable(changed between species) sites in a sample 
  • Pi: Based on the average number of differences between sequences. More sensitive to the frequency of a variation. 

Bioinformatics Week 6

Coursera Week 5: Phylogenetics
To study Phylogenetics you create visual comparisons between DNA sequences or proteins in the form of a tree. 
Mutation leads to speciation
 
There are rooted trees and unrooted trees. Rooted trees are the ones above, they expand from one point in one direction whereas unrooted trees can go off in many directions. Homeoplasy is when 2 divergent species share a similar characteristic. There are different types of homeoplasy. 
In order to conduct experiments involving phylogenetic you must have good sampling. Some of your samples need to be homologous, independent and variants of the original specimen which the tree is being based off of. Lastly, you need sequence alignments, and statistical support for their arrangement on the tree.

There are two tree building methods:
Distance methods

  • UPGMA
  • Neighbour Joining: Using blosum or PAM matrix to compare, then create a system to rate and scale the distance between the species based on their matrix score. 
  • Good things: They are computationally fast, and there is a singular best tree found in the end. 
  • Bad things: Sometimes there isn't a single best tree 

Character based (discrete) methods

  • Maximum parisomony
  • Maximum likelihood: Evaluates the likelihood of every possible mutation that could occur within a phylogenetic tree for a species to arrive at where it currently is. Then it uses statistical analysis to figure out which has the highest likelihood and assumes that's the correct tree. There are 4 base pairs so in an unbiased model there is a .25 likelihood for one of the 4 base pairs to change to another base pair. Then you multiply it to the 10th with the power of how many nucleotides there are within the sequence you are analyzing and that is the likelihood of a certain mutation. Say you have a sequence 20 base pairs long, and you can say that a certain G substitution you are studying has a .25*10^20 chance of occurring. Then after that it calculates the chances of the this change occurring over time in this fashion (process portion). Advantages: Produces clear results, you can statistically analyze the results you receive, it also gives you the other likely options that it produces. Disadvantages: It is computationally intensive and cannot be applied to large datasets. 
There is something called bootstrapping where you take all possible versions of your phylogenetic tree and then you calculate how many times certain species are grouped together. Here we see that A and B have been grouped together 100% (this number is arbitrary)  and C and D have been grouped together 75%. So it is very likely that A and B and then C and D diverge from a more recent ancestor. 70-90% the relationship is very probable. Anything less means it is a less probable relationship. 

Sunday, August 7, 2016

Bioinformatics Week 5

Coursera Week 3:
Multiple sequence alignments allow you to see the evolution of a species, as well as figuring out which sequences are useful through which sequences are preserved.
To do multiple sequence alignments you need to create a scoring guideline. You must compare columns and then assign a numerical value to rank to homologous columns.
Algorithms that code for MSA (multiple sequence alignment):
Dynamic (better)
Multidimensional dynamic (worse)
Programs that do MSA are Clustal (Progressive MSA) or DIALIGN (Local MSA)

Progressive MSA progressively aligns more distantly related sequences. The sequence is not disturbed during alignment.
Clustal
Insert gaps if you need to in order to better align the sequences
If you place a gap within a sequence you must also add in a deficit for the gap. (subtract points from alignment score for gap insertion.)
Then create a guide tree based on how related the sequences are
These guide trees are phylogenic trees
Clustal is suffers from making the quickest solution rather than the best solution. It looks for temporary fixes (inserting gaps in the quickest place rather than the most strategic) rather than long term fixes. These temporary fixes eventually propagate and lead to poorer total alignment. In order to compensate for the errors made by Clustal there are iterative methods that go through the alignments and then identify the subgroups within the larger allignments that have been aligned with quick fixes, then it fixes the temporary fixes with long term fixes and then reinsert them into the total sequence to be realigned.
Once Iterative programs have fixed the temporary fixes, it goes through and then fixes the alignments again, it then goes through the newly aligned sequence it has just created and creates a phylogenetic tree, then it determines the MSA, scores the MSA and then compares that score with the original score of the Clustal alignment it fixed, and asks whether the score is better or not. If its better it goes back and realigns everything again and if its not better then its done. It runs in a circle constantly trying to better the alignment, each time it gets a better alignment it keeps realigning till it hits a point when the alignment cannot get better.
Dynamic substitution matrices are used in order to compare sequences once they are aligned. It uses lower value Blosum matrices for lower scored alignments.  (These matrices are used in BLAST).

Then there is Local MSA (DIALIGN)
Which compares sequences of DNA within a global sequence (total sequence) which are in different places known as diagonals.
 This is basically what DIALIGN does. It then weighs the worth of each diagonal, based on length.

To compare sequences which are related use Clustal
To compare sequences which are unrelated and have conserved regions that are consistent use DIALIGN
To compare sequences which are unrelated and have conserved regions that are non consistent use MAFFT
Protein is easier to align than DNA. DNA gets the score of 1 if it matches, 0 if it doesn't. But proteins have amino acids which are redundant so it's easier to find a match.
*Too many caps or insertions or columns that don't match means that something is wrong with the alignment.
Using MSA programs is a skill. It is very easy to get poor MSA results using Clustal (the most popular website.) While using any of the MSA programs it is very important that it is not full proof and you cannot trust the results produced.
When you use global alignments:
When the sequences can be aligned through the entire sequence. If the sequences are of different lengths then you can insert gaps in order to compensate for the different sizes and then align.
When you use local alignments:
When the sequence can only be aligned at certain areas of the sequence.
When you use NCBI downloads of sequences in order to input them into MSA programs it is important to note that their names are incredibly lengthy. You must learn Perl, Python or Ruby in order to rename the files and make things very simple.
MEGA:
MSA program MEGA is very good at taking DNA translating it into protein, aligning it and then retranslating it back into DNA. This is a very useful tool, however it must be used carefully because only a small percentage of DNA sequences code for protein. You also must make sure that the sequence you are inputting is the full sequence, if you start at a different point than the starting point then the sequence is read incorrectly because it is read in codons and it will code for the wrong protein.
DIALIGN
When you have sequences that are unalignable but have short conserved regions, it is best to use DIALIGN. DIALIGN is also very good at the DNA to protein conversion.
MAFFT
The best tool to use for general MSA problems that MEGA or Clustal struggle with. The setbacks presented with the Clustal process are fixed with MAFFT. MAFFT automatically accounts for the quality of the MSA score by looking at the number of inserts and lessening the MSA score accordingly. It also works at incredible speeds considering the amount of work it is doing.
Programing words:
Regex (regular expressions) this is programing a system to associate what you input as what it has in its system. For example, looking up obvi and having obviously come up. It makes it easier and more efficient for you to research. You will have to know Ruby, Perl or Python to do so. Crimson is a good way to work with regex.
***MSA is integral for solving the issue the lyme disease paper brought up. The genomes of all of the strains of lyme disease must be sequenced and after they are sequenced you need to run them through an MSA program in order to see which parts of the sequence are conserved throughout the species and which are not. This can help determine which sequences are essential to the functioning of lyme disease and which sequences are associated with the differentiation and adaption of the strains. Because MSA uses the scoring Matrices it is also vital to master those.

Sunday, July 31, 2016

Summer Research Week 5 (08/01 - 08/05)

Great progress! This week there will be two major focuses:

Bioinformatics
  1. Bioinformatics Methods I, Coursera: Go through the materials of week 3.
Python & NLTK
  1. In order to apply Natural Language Processing (NLP) to biomedical fields, you will have to learn a programming language- Python and a platform- Natural Language Toolkit (NLTK). 
  2. Go to https://www.python.org/downloads/ to download and install the most recent version (3.5.2) of Python.   
  3. Launch the "Terminal" in the Applications > Utilities folder. In the terminal, run the following commands: (Let me know if you have any problem. I have tried a few times, and finally made it work.)
    • Install pip: run sudo easy_install pip 
    • Install NLTK: run sudo pip install -U nltk 
    • Install Numpy: run sudo pip install -U numpy 
    • Run Python: run python
    • Test installation: type import nltk
  4. Now, you can go to http://www.nltk.org/, and follow the "Some simple things you can do with NLTK".