Braineos
DiscoverChallenges
Sign in
← Back to decks
Sign in to save your progress, vote, and build your own decks.Sign in

CSB472-Final Exam

99 cards·by SGlucose
Study this deck
Needleman-Wunsch
dynamic programming global alignment
Smith-Waterman
dynamic programming local alignment
computational genomics
create ways to analysis the genome, build on the genome
bioinformatics
applies the techniques created by computational genomics
similarity
quantitative measure, can imply homology(primary way that bioinformatics infers relationships)
homology
qualitatve statment about common descent
orthology
duplication b/c of a speciation event
paralogy
duplicaiton b/c of duplication within an organism
Xenologous genes
related through a horizontal gene transfer event
synteny
describes the physical co-localization of genetic loci on the same chromosome within an individual or species
soft masking
only masks for the seeding step
blastn
NT(query) to NT-->NT
tblastx
NT(query) to NT -->AA
blastx
NT(query) to AA-->AA
tblastn
AA(query) to NT-->AA
blastp
AA to AA-->AA
E-Value
Expect Threshold-a parameter that describes the number of hits one can "expect" by chance when searching a data based of this size
bl2seq
local pairwise alignment
dot matrix alignment
see lect 2
Kmer
like a word hit, but for FAST
Bon-Ferony Correction
(p-value)/[number of this searched(BLAST significant) or number of genes analysed(microarays)]
alignment score
log-odds score (for a given amino acid pair)
Sum-of-Pairs Scoring
MSA scoring method
Clustal
Global and Progressive MSA program
muscle
a good iterative MSA
overlap wieght
how dialign diagonals are scored
edge
branch on a phylogenic map
nodes
where phylogenic trees branch
OTU
operational taxonomical unit
ancestral state
the state of the most rescent common ancestor
singltons
mutations that occur in only one end point
parallel evolution
independent evolution to same character from same ancestral state
coevolution
independent evolution to same character from different ancestral state
secondary loss
reversion to the ancesteral state
UPGMA
Unweighted Pair Grouping Method using Averages, a Distance Based Phylogentic Method
Neighbour Joining
Distance Based Phylogentic Method
regular expression
pattern
2 distance based phylogenetic methods
UPGMA and neighbor-joining
2 character-based phylogenetic methods
max parsimony and max likilihood
transitions
purine to purine or pyrimidine to pyrimidine
transversion
purine < --> pyrimidine
weighted parsimony
max parsimony phylogenetic method, but with weighting on different substituation
Kimura 2-Parameter
equal freq and unequal Tn vs Tv
Jukes-Cantor
calculates distance from substitutions (equal freq and equal rate)
F81
unequal freq and equal rates
KHY85
unequal freq and unequal Tn vs Tv
General Time-Reversal
unequal freq and six unequal rates
Rate Heterogeneity
refers to the idea that different sites will have different substitution rates
psuedoreplicates
re-sampled alignments (boot strapping)
tuple
a record in a relational database
attributes
fields or column in a relational database
most biological databased are what format
flat-file or relational
SLQ and its uses...
structured query language, used to create or query databases
XML, advantages/disadvantage...
extensible markup languange. it is NOT a flat file format. explicit tagging but larger file size.
Identifier
understandable by a human, less stable than an accession number
accession code/number
a unique code for a given entry, often arbitrary
RefSeq
provides stable reference sequences of all molec involve din the central dogma
UniGene
cluster EST and mRNA seq aong with conding seq and annotated genomic DNA in subsets of related seq
PROSITE
protien motif database
Pfam
protein motif database
ontology
a controlled vocabulary for describing a knowledge systmem
GEO
gene expression database
ArrayExpress
gene expression database
Entrez
for searching across databases
N50
The maximum length L such that 50% of all bases lie in contigs atleast L bases long. generally,the longer the better
species complexity
number of species
PRINTs
database of the fingerprints of different protein families
fingerprint
the set of motifs common to a protein family. the motifs are unweighted
shannon entropy
(sum the freq of a given animo acid in the entire allignemnt)* log2 (of the aminoacid at the specific position we are looking at)
profiles
, it is possible to create a statistical model of a protein family. AKA PSSMs
finite state machine
moves through a series of states and emits an output
Forward algorithum
a heuristic method used of HMMs
Viterbi
dynamic programin based heuristic approach to using HMMs
Pfam
a large collection of multiple sequence alignments and hidden Markov models covering many common protein families
Superfam
a large collections of HMMs that were created based on structural similarity
InterPro
database that integrates a bunch of different info on protein(structural, motif, HMM ect...)
max bit score
Information content. think about it as the number of binary question (yes/no) question you must ask to find out what is at a given position
metagenomics
using genomic techniques to analysis groups of microorganisms in thier natural enviroments
Fickett's TestCode
Looks at 3rd asymmetry in possition use
Uneven Positional Base Frequencies:
looks at NT usage biases in all three codon possitions A-DIF/T-DIF/ect..
Positional Base Preference
compairs the codons in all three reading frame to a prot of average AA comp, but not codon biases
Shine- Delgarno seq
prok transcription initiation seq
MA plot
log (ratio) vs log (intensity)
Pseudomedia
a measure of centrality for data-sets and populations. It agrees with the median for symmetric data-sets or populations.
MIAME
specify the metadata required for microarray submitions (e.g. tissue type and age and source)
MAGE-ML
specify the metadata required for microarray submitions. more comprehensive that MIAME
FWER
the probability of at least one fase positive in all the tests
FDR
fractions of tests that were FPs given some cut off value used to judge significance
SAM
Significance analysis of microarrays
Hypergeometric P-value
a way of analysis if there is enrichment in your gene set for a given catagory
protein microarrays
an array with proteins bound to te slide
glycan arrays
an array with sugar boud to the slide
HSP
High scoring segment pairs
Threshold Word Score
cut off score used when determining the extended word list in BLAST
lambda(for normalized bit score)
a factor that turns sequence score into a probability
K (for normalized bit scores)
a factor to account for the number of comparisions
converting E-values to P-values
P=1-e^-E
match model
P(x,y|M)
conditional substituation probability
used when making PAM, see lect 4