Tuesday, 18 July 2017

WHAT IS BIOINFORMATICS?
(Molecular) bio informatics: bioinformatics is conceptualising biology in terms of molecules (in the sense of physical chemistry) and applying "informatics techniques" (derived from disciplines such as applied maths, computer science and statistics) to understand and organise the information associated with these molecules, on a large scale. In short, bioinformatics is a management information system for molecular biology and has many practical applications.

Broadly speaking, Bioiformatics or computational biology is the application of computer science, statistics, and mathematics to problems in biology. Computational biology spans a wide range of fields within biology, including genomics/genetics, biophysics, cell biology, biochemistry, and evolution. Likewise, it makes use of tools and techniques from many different quantitative fields, including algorithm design, machine learning, Bayesian and frequentist statistics, and statistical physics.

What kinds of problems do computational biologists work on?

Much of computational biology is concerned with the analysis of molecular data, such as biosequences (DNA, RNA, or protein sequences), three-dimensional protein structures, gene expression data, or molecular biological networks (metabolic pathways, protein-protein interaction networks, or gene regulatory networks). A wide variety of problems can be addressed using these data, such as the identification of disease-causing genes, the reconstruction of the evolutionary histories of species, and the unlocking of the complex regulatory codes that turn genes on and off. Computational biology can also be concerned with non-molecular data, such as clinical or ecological data.

What are the differences between computational biology and bioinformatics?

The terms computational biology and bioinformatics are often used interchangeably. However, computational biology sometimes connotes the development of algorithms, mathematical models, and methods for statistical inference, while bioinformatics is more associated with the development of software tools, databases, and visualization methods.

For your Classwork No. 1

GACCTACACCTGTCAACATAATTGGAAGAAATCTGTTGACTCAGATTGGTTGCACTTTAAATTTTCCCATTAGCCCTATTGAGACTGTACCAGTAAAATTAAAGCCAGGAATGGATGGCCCAAAAGTTAAACAATGGCCATTGACAGAAGAAAAAATAAAAGCATTAGTAGAAATTTGTACAGAGATGGAAAAGGAAGGGAAAATTTCAAAAATTGGGCCTGAAAATCCATACAATACTCCAGTATTTGCCATAAAGAAAAAAGACAGTACTAAATGGAGAAAATTAGTAGATTTCAGAGAACTTAATAAGAGAACTCAAGACTTCTGGGAAGTTCAATTAGGAATACCACATCCCGCAGGGTTAAAAAAGAAAAAATCAGTAACAGTACTGGATGTGGGTGATGCATATTTTTCAGTTCCCTTAGATGAAGACTTCAGGAAGTATACTGCATTTACCATACCTAGTATAAACAATGAGACACCAGGGATTAGATATCAGTACAATGTGCTTCCACAGGGATGGAAAGGATCACCAGCAATATTCCAAAGTAGCATGACAAAAATCTTAGAGCCTTTTAGAAAACAAAATCCAGACATAGTTATCTATCAATACATGGATGATTTGTATGTAGGATCTGACTTAGAAATAGGGCAGCATAGAACAAAAATAGAGGAGCTGAGACAACATCTGTTGAGGTGGGGACTTACCACACCAGACAAAAAACATCAGAAAGAACCTCCATTCCTTTGGATGGGTTATGAACTCCATCCTGATAAATGGACAGTACAGCCTATAGTGCTGCCAGAAAAAGACAGCTGGACTGTCAATGACATACAGA


1. What is the name of the protein that contain this DNA sequence?
2. In which organism does was this sequence derived from?
3. When was the sequence loaded in the database?
4. What is the Accession number of the sequence?
5. What is the percentage identity of the sequence to the query seuence?
6. What is the Expectation(E)-value of the results?

Monday, 17 July 2017

Sequence Databases

The NCBI Sequence Database

All published genome sequences are available over the internet, as it is a requirement of every scientific journal that any published DNA or RNA or protein sequence must be deposited in a public database. The main resources for storing and distributing sequence data are three large databases: the NCBI database (www.ncbi.nlm.nih.gov/), the European Molecular Biology Laboratory (EMBL) database (www.ebi.ac.uk/embl/, and the DNA Database of Japan (DDBJ) database (www.ddbj.nig.ac.jp/). These databases collect all publicly available DNA, RNA and protein sequence data and make it available for free. They exchange data nightly, so contain essentially the same data.

In this chapter we will discuss the NCBI database. Note however that it contains essentially the same data as in the EMBL/DDBJ databases.
Sequences in the NCBI Sequence Database (or EMBL/DDBJ) are identified by an accession number. This is a unique number that is only associated with one sequence. For example, the accession number NC_001477 is for the DEN-1 Dengue virus genome sequence. The accession number is what identifies the sequence. It is reported in scientific papers describing that sequence.
As well as the sequence itself, for each sequence the NCBI database (or EMBL/DDBJ databases) also stores some additional annotation data, such as the name of the species it comes from, references to publications describing that sequence, etc. Some of this annotation data was added by the person who sequenced a sequence and submitted it to the NCBI database, while some may have been added later by a human curator working for NCBI.
The NCBI database contains several sub-databases, the most important of which are:
  • the NCBI Nucleotide database: contains DNA and RNA sequences
  • the NCBI Protein database: contains protein sequences
  • EST: contains ESTs (expressed sequence tags), which are short sequences derived from mRNAs
  • the NCBI Genome database: contains DNA sequences for whole genomes
  • PubMed: contains data on scientific publications
Classwork 2

Q1. What information about the rabies virus sequence (NCBI accession NC_001542) can you obtain from its annotations in the NCBI Sequence Database?
What does it say in the DEFINITION and ORGANISM fields of its NCBI record? Note: rabies virus is the virus responsible for rabies, which is classified by the WHO as a neglected tropical disease.
Q2. How many nucleotide sequences are there from the bacterium Chlamydia trachomatis in the NCBI Sequence Database?
Note: the bacterium Chlamydia trachomatis is responsible for causing trachoma, which is classified by the WHO as a neglected tropical disease.