viernes, 20 de noviembre de 2015

Statistical Tests for Clonality

Statistical Tests for Clonality

SUMMARY

Cancer investigators frequently conduct studies to examine tumor samples from pairs of apparently independent primary tumors with a view to determining if they share a “clonal” origin. The genetic fingerprints of the tumors are compared using a panel of markers, often representing loss of heterogeneity (LOH) at distinct genetic loci. In this article we evaluate candidate significance tests for this purpose. The relevant information derives from the observed correlation of the tumors with respect to the occurrence of LOH at individual loci, a phenomenon that can be evaluated using Fisher’s Exact Test. Information is also available from the extent to which losses at the same locus occur on the same parental allele. Data from these combined sources of information can be evaluated using a simple adaptation of Fisher’s Exact Test. The test statistic is the total number of loci at which concordant mutations occur on the same parental allele, with higher values providing more evidence in favor of a clonal origin for the two tumors. The test is shown to have high power for detecting clonality for plausible models of the alternative (clonal) hypothesis, and for reasonable numbers of informative loci, preferably located on distinct chromosomal arms. The method is illustrated using studies to identify clonality in contralateral breast cancer. Interpretation of the results of these tests requires caution due to simplifying assumptions regarding the possible variability in mutation probabilities between loci, and possible imbalances in the mutation probabilities between parental alleles. Nonetheless, we conclude that the method represents a simple, powerful strategy for distinguishing independent tumors from those of clonal origin.
Keywords: Clonality, Permutation test, Second primary cancers

Clonality: A Package for Clonality testing
Statistical Challenges in Testing Clonal





Molecular Evolution

Inference of Population Splits and Mixtures from Genome-Wide Allele Frequency Data

Abstract

Many aspects of the historical relationships between populations in a species are reflected in genetic data. Inferring these relationships from genetic data, however, remains a challenging task. In this paper, we present a statistical model for inferring the patterns of population splits and mixtures in multiple populations. In our model, the sampled populations in a species are related to their common ancestor through a graph of ancestral populations. Using genome-wide allele frequency data and a Gaussian approximation to genetic drift, we infer the structure of this graph. We applied this method to a set of 55 human populations and a set of 82 dog breeds and wild canids. In both species, we show that a simple bifurcating tree does not fully describe the data; in contrast, we infer many migration events. While some of the migration events that we find have been detected previously, many have not. For example, in the human data, we infer that Cambodians trace approximately 16% of their ancestry to a population ancestral to other extant East Asian populations. In the dog data, we infer that both the boxer and basenji trace a considerable fraction of their ancestry (9% and 25%, respectively) to wolves subsequent to domestication and that East Asian toy breeds (the Shih Tzu and the Pekingese) result from admixture between modern toy breeds and “ancient” Asian breeds. Software implementing the model described here, called TreeMix, is available at http://treemix.googlecode.com.

Abstract

Phylogenies of highly genetically variable viruses such as HIV-1 are potentially informative of epidemiological dynamics. Several studies have demonstrated the presence of clusters of highly related HIV-1 sequences, particularly among recently HIV-infected individuals, which have been used to argue for a high transmission rate during acute infection. Using a large set of HIV-1 subtype B pol sequences collected from men who have sex with men, we demonstrate that virus from recent infections tend to be phylogenetically clustered at a greater rate than virus from patients with chronic infection (‘excess clustering’) and also tend to cluster with other recent HIV infections rather than chronic, established infections (‘excess co-clustering’), consistent with previous reports. To determine the role that a higher infectivity during acute infection may play in excess clustering and co-clustering, we developed a simple model of HIV infection that incorporates an early period of intensified transmission, and explicitly considers the dynamics of phylogenetic clusters alongside the dynamics of acute and chronic infected cases. We explored the potential for clustering statistics to be used for inference of acute stage transmission rates and found that no single statistic explains very much variance in parameters controlling acute stage transmission rates. We demonstrate that high transmission rates during the acute stage is not the main cause of excess clustering of virus from patients with early/acute infection compared to chronic infection, which may simply reflect the shorter time since transmission in acute infection. Higher transmission during acute infection can result in excess co-clustering of sequences, while the extent of clustering observed is most sensitive to the fraction of infections sampled.

A general linear model-based approach for inferring selection to climate


Estimation of Population Genetic Structure Software and Methods

Artìculos útiles para Estimación de Estructura poblacional:

On Identifying the Optimal Number of Population Clusters via the Deviance Information Criterion
On Identifying the... 

Detecting correlation between allele frequencies and environmental variables as a signature of selection. A fast computational approach for genome-wide studies

Detecting and measuring selection from gene frequency data












GenClone 2.0

GenClone: a computer program to analyze genotypic data, test for clonality and describe spatial clonal organization
Arnaud-Haond Sophie and Belkhir Khalid
«Team MAREE» - CCMAR, Algarve University, FCMA, Gambelas, 8005-139 Faro, PORTUGAL
«Génome, Populations, Interactions »-Université Montpellier II, Place Eugène Bataillon ; 34090 Montpellier Cedex, FRANCE

Link: GenClone

lunes, 4 de julio de 2011

Novel Phylogenetic Inference Software!

 http://www.metapiga.org/welcome.html
MetaPIGA 2 is a robust implementation of several stochastic heuristics for large phylogeny inference (under maximum likelihood), including a random-restart hill climbing, a simulated annealing algorithm, a classical genetic algorithm, and the metapopulation genetic algorithm (metaGA) together with complex substitution models, discrete Gamma rate heterogeneity, and the possibility to partition data. MetaPIGA 2 handles nucleic-acid and protein datasets as well as morphological (presence/absence) data. The benefits of the metaGA (Lemmon & Milinkovitch 2002; PNAS, 99: 10516-10521) are as follows: (i) it resolves the major problem inherent to classical Genetic Algorithms (i.e., the need to choose between strong selection, hence, speed, and weak selection, hence, accuracy) by maintaining high inter-population variation even under strong intra-population selection, and (ii) it can generate branch support values that approximate posterior probabilities.
The software MetaPIGA 2 also implements:

  • Simple dataset quality control (testing for the presence of identical sequences as well as for excessively ambiguous or excessively divergent sequences);
  • Automated trimming of poorly aligned regions using the trimAl algorithm;
  • The Likelihood Ratio Test, the Akaike Information Criterion, and the Bayesian Information Criterion for the easy selection of nucleotide and amino-acid substitution models that best fit the data;
  • Ancestral-state reconstruction of all nodes in the tree.
MetaPIGA 2 provides high customization of heuristics' and models' parameters, manual batch file and command line processing. However, it also offers an extensive and ergonomic graphical user interface and functionalities assisting the user for dataset quality testing, parameters setting, generating and running batch files, following run progress, and manipulating result trees.
MetaPIGA 2 uses standard formats for data sets and trees, is platform independent, runs in 32- and 64-bits systems, and takes advantage of multiprocessor and/or multicore computers. A version for Grid computing is in development.
 

Citing MetaPIGA 2

MetaPIGA v2.0: maximum likelihood large phylogeny estimation using the metapopulation genetic algorithm and other stochastic heuristics
Raphaël Helaers & Michel C. Milinkovitch
BMC Bioinformatics 2010, 11:379



http://bioinformatics.oxfordjournals.org/content/25/2/197.full

Phylogenetic inference under recombination using Bayesian stochastic topology selection


Abstract

Motivation: Conventional phylogenetic analysis for characterizing the relatedness between taxa typically assumes that a single relationship exists between species at every site along the genome. This assumption fails to take into account recombination which is a fundamental process for generating diversity and can lead to spurious results. Recombination induces a localized phylogenetic structure which may vary along the genome. Here, we generalize a hidden Markov model (HMM) to infer changes in phylogeny along multiple sequence alignments while accounting for rate heterogeneity; the hidden states refer to the unobserved phylogenic topology underlying the relatedness at a genomic location. The dimensionality of the number of hidden states (topologies) and their structure are random (not known a priori) and are sampled using Markov chain Monte Carlo algorithms. The HMM structure allows us to analytically integrate out over all possible changepoints in topologies as well as all the unknown branch lengths.

Results: We demonstrate our approach on simulated data and also to the genome of a suspected HIV recombinant strain as well as to an investigation of recombination in the sequences of 15 laboratory mouse strains sequenced by Perlegen Sciences. Our findings indicate that our method allows us to distinguish between rate heterogeneity and variation in phylogeny caused by recombination without being restricted to 4-taxa data.

Availability: The method has been implemented in JAVA and is available, along with data studied here, from http://www.stats.ox.ac.uk/~webb.

Contact: cholmes@stats.ox.ac.uk

Supplementary information: Supplementary data are available at Bioinformatics online.



http://www.stats.ox.ac.uk/__data/assets/pdf_file/0005/4010/large_pedigrees.pdf


http://www.cs.cmu.edu/~guestrin/Class/10701-S07/Handouts/recitations/HMM-inference.pdf


Probabilistic Phylogenetic Inference with Insertions and Deletions

Abstract Top

A fundamental task in sequence analysis is to calculate the probability of a multiple alignment given a phylogenetic tree relating the sequences and an evolutionary model describing how sequences change over time. However, the most widely used phylogenetic models only account for residue substitution events. We describe a probabilistic model of a multiple sequence alignment that accounts for insertion and deletion events in addition to substitutions, given a phylogenetic tree, using a rate matrix augmented by the gap character. Starting from a continuous Markov process, we construct a non-reversible generative (birth–death) evolutionary model for insertions and deletions. The model assumes that insertion and deletion events occur one residue at a time. We apply this model to phylogenetic tree inference by extending the program DNAML in PHYLIP. Using standard benchmarking methods on simulated data and a new "concordance test" benchmark on real ribosomal RNA alignments, we show that the extended program DNAMLε improves accuracy relative to the usual approach of ignoring gaps, while retaining the computational efficiency of the Felsenstein peeling algorithm.

Author Summary Top

We describe a computationally efficient method to use insertion and deletion events, in addition to substitutions, in phylogenetic inference. To date, many evolutionary models in probabilistic phylogenetic inference methods have only accounted for substitution events, not for insertions and deletions. As a result, not only do tree inference methods use less sequence information than they could, but also it has remained difficult to integrate phylogenetic modeling into sequence alignment methods (such as profiles and profile-hidden Markov models) that inherently require a model of insertion and deletion events. Therefore an important goal in the field has been to develop tractable evolutionary models of insertion/deletion events over time of sufficient accuracy to increase the resolution of phylogenetic inference methods and to increase the power of profile-based sequence homology searches. Our model offers a partial answer to this problem. We show that our model generally improves inference power in both simulated and real data and that it is easily implemented in the framework of standard inference packages with little effect on computational efficiency (we extended DNAML, in Felsenstein's popular PHYLIP package).



Materials and Methods Top

The C source code for the modified PHYLIP 3.66 package [14] that contains the program DNAMLε , the C source code for evolving sequences with the generative model (εRATE ), the modified ROSE package (version 1.3) [76], as well as all the Perl scripts and datasets used to generate the results presented in this paper are provided as a tarball in Dataset S1. The program DNAMLε uses the EASEL sequence analysis library (SRE, unpublished) which is also provided.

Roland F. Schwarz, William Fletcher, Frank Förster, Benjamin Merget, Matthias Wolf, Jörg Schultz, and Florian Markowetz
PLoS One. 2010; 5(12): e15788. Published online 2010 December 31. doi: 10.1371/journal.pone.0015788
PMCID:
PMC3013127

Bhakti Dwivedi and Sudhindra R Gadagkar
BMC Evol Biol. 2009; 9: 211. Published online 2009 August 23. doi: 10.1186/1471-2148-9-211
PMCID:
PMC2746219


Title: A stochastic evolution model for residue Insertion-Deletion Independent from Substitution
Author(s): Lebre S, Michel CJ
Source: COMPUTATIONAL BIOLOGY AND CHEMISTRY   Volume: 34   Issue: 5-6   Pages: 259-267   Published: DEC 2010
Times Cited: 0

Title: Genomes as documents of evolutionary history
Author(s): Boussau B, Daubin V
Source: TRENDS IN ECOLOGY & EVOLUTION   Volume: 25   Issue: 4   Pages: 224-232   Published: APR 2010
Times Cited: 2


http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2746219/?tool=pmcentrez
Phylogenetic inference under varying proportions of indel-induced alignment gaps
Bhakti Dwivedi1 and Sudhindra R Gadagkarcorresponding author1,2
1Department of Biology, University of Dayton, 300 College Park, Dayton, OH 46469-2320, USA
2Department of Natural Sciences, PO Box 1004, 1400 Brush Row Rd, Wilberforce, Ohio 45384, USA
corresponding authorCorresponding author.
Bhakti Dwivedi: dwivedbz@notes.udayton.edu; Sudhindra R Gadagkar: sgadagkar@centralstate.edu
Received May 11, 2009; Accepted August 23, 2009.
Background
The effect of alignment gaps on phylogenetic accuracy has been the subject of numerous studies. In this study, we investigated the relationship between the total number of gapped sites and phylogenetic accuracy, when the gaps were introduced (by means of computer simulation) to reflect indel (insertion/deletion) events during the evolution of DNA sequences. The resulting (true) alignments were subjected to commonly used gap treatment and phylogenetic inference methods.
Results
(1) In general, there was a strong – almost deterministic – relationship between the amount of gap in the data and the level of phylogenetic accuracy when the alignments were very "gappy", (2) gaps resulting from deletions (as opposed to insertions) contributed more to the inaccuracy of phylogenetic inference, (3) the probabilistic methods (Bayesian, PhyML & "MLε, " a method implemented in DNAML in PHYLIP) performed better at most levels of gap percentage when compared to parsimony (MP) and distance (NJ) methods, with Bayesian analysis being clearly the best, (4) methods that treat gapped sites as missing data yielded less accurate trees when compared to those that attribute phylogenetic signal to the gapped sites (by coding them as binary character data – presence/absence, or as in the MLε method), and (5) in general, the accuracy of phylogenetic inference depended upon the amount of available data when the gaps resulted from mainly deletion events, and the amount of missing data when insertion events were equally likely to have caused the alignment gaps.
Conclusion
When gaps in an alignment are a consequence of indel events in the evolution of the sequences, the accuracy of phylogenetic analysis is likely to improve if: (1) alignment gaps are categorized as arising from insertion events or deletion events and then treated separately in the analysis, (2) the evolutionary signal provided by indels is harnessed in the phylogenetic analysis, and (3) methods that utilize the phylogenetic signal in indels are developed for distance methods too. When the true homology is known and the amount of gaps is 20 percent of the alignment length or less, the methods used in this study are likely to yield trees with 90–100 percent accuracy.
 
PICS-Ord: unlimited coding of ambiguous regions by pairwise identity and cost scores ordination
Robert Lücking, Brendan P Hodkinson, Alexandros Stamatakis, and Reed A Cartwright
BMC Bioinformatics. 2011; 12: 10. Published online 2011 January 7. doi: 10.1186/1471-2105-12-10.
PMCID: PMC3024941
Phylogenetic assessment of alignments reveals neglected tree signal in gaps
Christophe Dessimoz and Manuel Gil
Genome Biol. 2010; 11(4): R37. Published online 2010 April 6. doi: 10.1186/gb-2010-11-4-r37.
PMCID: PMC2884540
| Abstract | Full Text | PDF–741K | Supplementary Material |



Stud Health Technol Inform. 2007;129(Pt 2):1245-9.

Enhancing the quality of phylogenetic analysis using fuzzy hidden Markov model alignments.

Source

Lab of Medical Informatics, Faculty of Medicine, Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki, Greece.

Abstract

Any effective phylogeny inference based on molecular data begins by performing efficient multiple sequence alignments. So far, the Hidden Markov Model (HMM) method for multiple sequence alignment has been proved competitive to the classical deterministic algorithms with respect to phylogenetic analysis; nevertheless, its stochastic nature does not help it cope with the existing dependence among the sequence elements. This paper deals with phylogenetic analysis of protein and gene data using multiple sequence alignments produced by fuzzy profile Hidden Markov Models. Fuzzy profile HMMs are a novel type of profile HMMs based on fuzzy sets and fuzzy integrals, which generalize the classical stochastic HMM by relaxing its independence assumptions. In this paper, alignments produced by the fuzzy HMM model are used in phylogenetic analysis of protein data, enhancing the quality of phylogenetic trees. The new methodology is implemented in HPV virus phylogenetic inference. The results of the analysis are compared against those obtained by the classical profile HMM model and depict the superiority of the fuzzy profile HMM in this field.


Bioinformatics. 2005 Sep 1;21 Suppl 2:ii166-72.

Discriminating between rate heterogeneity and interspecific recombination in DNA sequence alignments with phylogenetic factorial hidden Markov models.

Source

Biomathematics and Statistics, Scotland, Edinburgh, UK. dirk@bioss.ac.uk

Abstract

MOTIVATION:

A recently proposed method for detecting recombination in DNA sequence alignments is based on the combination of hidden Markov models (HMMs) with phylogenetic trees. Although this method was found to detect breakpoints of recombinant regions more accurately than most existing techniques, it inherently fails to distinguish between recombination and rate variation. In the present paper, we propose to marry the phylogenetic tree to a factorial HMM (FHMM). The states of the first hidden chain represent tree topologies, whereas the states of the second independent hidden chain represent different global scaling factors of the branch lengths. Inference is done in terms of a hierarchical Bayesian model, where parameters and hidden states are sampled from the posterior distribution with Gibbs sampling.

RESULTS:

We have tested the proposed model on various synthetic and real-world DNA sequence alignments. The simulation results suggest that as opposed to the standard phylogenetic HMM, the phylogenetic FHMM clearly distinguishes between recombination and rate heterogeneity and thereby avoids the prediction of spurious recombinant regions.

AVAILABILITY:

The proposed method has been implemented in a MATLAB package that extends Kevin Murphy's HMM toolbox. Software and data used in our study are available from http://www.bioss.sari.ac.uk/~dirk/Supplements


martes, 1 de diciembre de 2009


jueves, 9 de julio de 2009

allele determination

http://crop.scijournals.org/cgi/content/full/46/5/2084
Crop Science PLANT GENETIC RESOURCES
Accuracy and Reliability of High-Throughput Microsatellite Genotyping for Cacao Clone Identification


http://www-naweb.iaea.org/nafa/aph/stories/dna-manual.pdf
A practical approach to microsatellite genotyping with special reference to livestock population genetics

allele determination

http://www.plantmethods.com/content/1/1/3
High throughput, high resolution selection of polymorphic microsatellite loci for multiplex analysis


http://www.gse-journal.org/index.php?option=article&access=standard&Itemid=129&url=/articles/gse/pdf/2002/03/g340301.pdf

http://www.cdfd.org.in/jnagpdf/jfc.pdf
Capillary Electrophoresis Is Essential for Microsatellite Marker Based Detection and Quantification of Adulteration of Basmati Rice (Oryza sativa)

http://crop.scijournals.org/cgi/content/full/43/5/1828
A Low-Cost, High-Throughput Polyacrylamide Gel Electrophoresis System for Genotyping with Microsatellite DNA Markers

allele age dtermination

P. Fearnhead, R. M. Harding, J. A. Schneider, S. Myers, and P. Donnelly
Application of Coalescent Methods to Reveal Fine-Scale Rate Variation and Recombination Hotspots
Genetics, August 1, 2004; 167(4): 2067 - 2081.


J. M. Burrows, L. Bromham, M. Woolfit, G. Piganeau, J. Tellam, G. Connolly, N. Webb, L. Poulsen, L. Cooper, S. R. Burrows, et al.
Selection Pressure-Driven Evolution of the Epstein-Barr Virus-Encoded Oncogene LMP1 in Virus Isolates from Southeast Asia
J. Virol., July 1, 2004; 78(13): 7131 - 7137.


P. Lemey, O. G. Pybus, A. Rambaut, A. J. Drummond, D. L. Robertson, P. Roques, M. Worobey, and A.-M. Vandamme
The Molecular Population Genetics of HIV-1 Group O
Genetics, July 1, 2004; 167(3): 1059 - 1068.


J. D. Wall
Estimating Recombination Rates Using Three-Site Likelihoods
Genetics, July 1, 2004; 167(3): 1461 - 1473.


D. T. Haydon, A. D. S. Bastos, and P. Awadalla
Low linkage disequilibrium indicative of recombination in foot-and-mouth disease virus gene sequence alignments
J. Gen. Virol., May 1, 2004; 85(5): 1095 - 1100.


N. Li and M. Stephens
Modeling Linkage Disequilibrium and Identifying Recombination Hotspots Using Single-Nucleotide Polymorphism Data
Genetics, December 1, 2003; 165(4): 2213 - 2233.


S. D. Polley, W. Chokejindachai, and D. J. Conway
Allele Frequency-Based Analyses Robustly Map Sequence Sites Under Balancing Selection in a Malaria Vaccine Candidate Antigen
Genetics, October 1, 2003; 165(2): 555 - 561.


M. Anisimova, R. Nielsen, and Z. Yang
Effect of Recombination on the Accuracy of the Likelihood Method for Detecting Positive Selection at Amino Acid Sites
Genetics, July 1, 2003; 164(3): 1229 - 1236.

Microsatellite analysis and other Tools

Genetics - J.-F. Lefebvre and D. Labuda
Fraction of Informative Recombinations: A Heuristic Approach to Analyze Recombination Rates
Genetics, April 1, 2008; 178(4): 2069 - 2079.


Proceedings B
A. J McCarthy, M.-A. Shaw, and S. J Goodman
Pathogen evolution and disease emergence in carnivores
Proc R Soc B, December 22, 2007; 274(1629): 3165 - 3174.


Genetics
S. De Mita, J. Ronfort, H. I. McKhann, C. Poncet, R. El Malki, and T. Bataillon
Investigation of the Demographic and Selective Forces Shaping the Nucleotide Diversity of Genes Involved in Nod Factor Signaling in Medicago truncatula
Genetics, December 1, 2007; 177(4): 2123 - 2133.
-
D. Garrigan, S. B. Kingan, M. M. Pilkington, J. A. Wilder, M. P. Cox, H. Soodyall, B. Strassmann, G. Destro-Bisol, P. de Knijff, A. Novelletto, et al.
Inferring Human Population Sizes, Divergence Times and Rates of Gene Flow From Mitochondrial, X and Y Chromosome Resequencing Data
Genetics, December 1, 2007; 177(4): 2195 - 2207.
-
J. Gay, S. Myers, and G. McVean
Estimating Meiotic Gene Conversion Rates From Population Genetic Data
Genetics, October 1, 2007; 177(2): 881 - 894.


D. T. Gerrard and A. Meyer
Positive Selection and Gene Conversion in SPP120, a Fertilization-Related Gene, during the East African Cichlid Fish Radiation
Mol. Biol. Evol., October 1, 2007; 24(10): 2286 - 2297.


S. R. Miller, R. W. Castenholz, and D. Pedersen
Phylogeography of the Thermophilic Cyanobacterium Mastigocladus laminosus
Appl. Envir. Microbiol., August 1, 2007; 73(15): 4751 - 4759.


A. Auton and G. McVean
Recombination rate estimation in the presence of hotspots
Genome Res., August 1, 2007; 17(8): 1219 - 1227.



A. RoyChoudhury and M. Stephens
Fast and Accurate Estimation of the Population-Scaled Mutation Rate, {theta}, From Microsatellite Genotype Data
Genetics, June 1, 2007; 176(2): 1363 - 1366.


X. Didelot and D. Falush
Inference of Bacterial Microevolution Using Multilocus Sequence Data
Genetics, March 1, 2007; 175(3): 1251 - 1266.


A. Ojeda, J. Rozas, J. M. Folch, and M. Perez-Enciso
Unexpected High Polymorphism at the FABP4 Gene Unveils a Complex History for Pig Populations
Genetics, December 1, 2006; 174(4): 2119 - 2127.


P. Fearnhead
SequenceLDhot: detecting recombination hotspots
Bioinformatics, December 15, 2006; 22(24): 3061 - 3066.


X. Liu, M. M. Gutacker, J. M. Musser, and Y.-X. Fu
Evidence for Recombination in Mycobacterium tuberculosis
J. Bacteriol., December 1, 2006; 188(23): 8169 - 8177.



D. S. Guttman, S. J. Gropp, R. L. Morgan, and P. W. Wang
Diversifying Selection Drives the Evolution of the Type III Secretion System Pilus of Pseudomonas syringae
Mol. Biol. Evol., December 1, 2006; 23(12): 2342 - 2354.


C. T. T. Edwards, E. C. Holmes, O. G. Pybus, D. J. Wilson, R. P. Viscidi, E. J. Abrams, R. E. Phillips, and A. J. Drummond
Evolution of the Human Immunodeficiency Virus Envelope Gene Is Dominated by Purifying Selection
Genetics, November 1, 2006; 174(3): 1441 - 1453.


S. L. Kosakovsky Pond, D. Posada, M. B. Gravenor, C. H. Woelk, and S. D. W. Frost
Automated Phylogenetic Detection of Recombination Using a Genetic Algorithm
Mol. Biol. Evol., October 1, 2006; 23(10): 1891 - 1901.


B. C. Verrelli, S. A. Tishkoff, A. C. Stone, and J. W. Touchman
Contrasting Histories of G6PD Molecular Evolution and Malarial Resistance in Humans and Chimpanzees
Mol. Biol. Evol., August 1, 2006; 23(8): 1592 - 1601.



P. L. Morrell, D. M. Toleno, K. E. Lundy, and M. T. Clegg
Estimating the Contribution of Mutation, Recombination and Gene Conversion in the Generation of Haplotypic Diversity
Genetics, July 1, 2006; 173(3): 1705 - 1723.


T. C. Bruen, H. Philippe, and D. Bryant
A Simple and Robust Statistical Test for Detecting the Presence of Recombination
Genetics, April 1, 2006; 172(4): 2665 - 2681.



A. Carvajal-Rodriguez, K. A. Crandall, and D. Posada
Recombination Estimation Under Complex Evolutionary Models with the Coalescent Composite-Likelihood Method
Mol. Biol. Evol., April 1, 2006; 23(4): 817 - 827.


C. Charpentier, T. Nora, O. Tenaillon, F. Clavel, and A. J. Hance
Extensive Recombination among Human Immunodeficiency Virus Type 1 Quasispecies Makes an Important Contribution to Viral Diversity in Individual Patients
J. Virol., March 1, 2006; 80(5): 2472 - 2482.


D. J. Wilson and G. McVean
Estimating Diversifying Selection and Functional Constraint in the Presence of Recombination
Genetics, March 1, 2006; 172(3): 1411 - 1425.


D. H. Bos and B. Waldman
Evolution by Recombination and Transspecies Polymorphism in the MHC Class I Gene of Xenopus laevis
Mol. Biol. Evol., January 1, 2006; 23(1): 137 - 143.


N. G. C. Smith and P. Fearnhead
A Comparison of Three Estimators of the Population-Scaled Recombination Rate: Accuracy and Robustness
Genetics, December 1, 2005; 171(4): 2051 - 2062.


K. Roselius, W. Stephan, and T. Stadler
The Relationship of Nucleotide Polymorphism, Recombination Rate and Selection in Wild Tomato Species
Genetics, October 1, 2005; 171(2): 753 - 763.


G. A.T McVean and N. J Cardin
Approximating the coalescent with recombination
Phil Trans R Soc B, July 29, 2005; 360(1459): 1387 - 1393.


L. Zhu and C. D. Bustamante
A Composite-Likelihood Approach for Detecting Directional Selection From DNA Sequence Data
Genetics, July 1, 2005; 170(3): 1411 - 1421.


D. Shriner, A. G. Rodrigo, D. C. Nickle, and J. I. Mullins
Pervasive Genomic Recombination of HIV-1 in Vivo
Genetics, August 1, 2004; 167(4): 1573 - 1583.

Microsatellite Allele Sizing and Analysis

doi:10.1016/S0021-9673(97)00542-6
Journal of Chromatography A
Volume 781, Issues 1-2, 26 September 1997, Pages 295-305
9th International Symposium on High Performance Capillary Electrophoresis and Related Microsale Techniques

Copyright © 1997 Published by Elsevier Science B.V.
Nucleic acids and their constituents
Rapid sizing of polymorphic microsatellite markers by capillary array electrophoresis

http://www.cdfd.org.in/jnagpdf/jfc.pdf
Capillary Electrophoresis Is Essential for Microsatellite Marker Based Detection and Quantification of Adulteration of Basmati Rice (Oryza sativa)


http://www.scfbm.org/content/3/1/15
Permutation – based statistical tests for multiple hypotheses

http://www3.interscience.wiley.com/journal/94516676/abstract?CRETRY=1&SRETRY=0
Permutation tests for detecting and estimating mixtures in task performance within groups

http://www.biomedcentral.com/1471-2105/9/511
Comparison of methods for estimating the nucleotide substitution matrix


http://www.genetics.org/cgi/content/abstract/160/3/1231
A Coalescent-Based Method for Detecting and Estimating Recombination From Gene Sequences



Genetics
M. Carneiro, N. Ferrand, and M. W. Nachman
Recombination and Speciation: Loci Near Centromeres Are More Differentiated Than Loci Near Telomeres Between Subspecies of the European Rabbit (Oryctolagus cuniculus)
Genetics, February 1, 2009; 181(2): 593 - 606.
[Abstract] [Full Text] [PDF]


Philosophical Transactions B
Y. Wang and B. Rannala
Bayesian inference of fine-scale recombination rates using population genomic data
Phil Trans R Soc B, December 27, 2008; 363(1512): 3921 - 3930.
[Abstract] [Full Text] [PDF]



Molecular Biology and Evolution
B. C. Verrelli, C. M. Lewis Jr, A. C. Stone, and G. H. Perry
Different Selective Pressures Shape the Molecular Evolution of Color Vision in Chimpanzee and Human Populations
Mol. Biol. Evol., December 1, 2008; 25(12): 2735 - 2743.
[Abstract] [Full Text] [PDF]



Genome Research
A. F. McRae, E. M. Byrne, Z. Z. Zhao, G. W. Montgomery, and P. M. Visscher
Power and SNP tagging in whole mitochondrial genome association studies
Genome Res., June 1, 2008; 18(6): 911 - 917.
[Abstract] [Full Text] [PDF]


doi:
10.1101/gr.8.1.69
Genome Res. 1998. 8: 69-80
Copyright © 1998, by Cold Spring Harbor Laboratory Press
High-Precision Genotyping by Denaturing Capillary Electrophoresis
H.-Michael Wenz, James M. Robertson, Steve Menchen, Frank Oaks, David M. Demorest, Don Scheibler, Barnett B. Rosenblum, Carla Wike, Dennis A. Gilbert, and J. William Efcavitch



From the Cover: Single-nucleotide polymorphism discovery by targeted DNA photocleavage
J. R. Hart, M. D. Johnson, and J. K. Barton
Proc. Natl. Acad. Sci. USA September 28, 2004 101: 14040-14044



Single nucleotide polymorphism detection by combinatorial fluorescence energy transfer tags and biotinylated dideoxynucleotides
A. K. Tong andJ. Ju
Nucleic Acids Res March 1, 2002 30: e19


Analysis of short tandem repeat polymorphisms by electrospray ion trap mass spectrometry
S. Hahner, A. Schneider, A. Ingendoh, and J. Mosner
Nucleic Acids Res September 15, 2000 28: e82

Highthrough-out put microsatellite allele sizing with high resolution





http://crop.scijournals.org/cgi/reprint/43/5/1828
Published in Crop Sci. 43:1828-1832 (2003).
© 2003 Crop Science Society of America
677 S. Segoe Rd., Madison, WI 53711 USA

GENOMICS, MOLECULAR GENETICS & BIOTECHNOLOGY
A Low-Cost, High-Throughput Polyacrylamide Gel Electrophoresis System for Genotyping with Microsatellite DNA Markers

D. Wang, J. Shi, S. R. Carlson, P. B. Cregan, R. W. Ward, and B. W. Diers*
Microsatellite DNA markers are widely used in genetic research. Their use, however, can be costly and throughput is sometimes limited. The objective of this paper is to introduce a simple, low-cost, high-throughput system that detects amplification products from microsatellite markers by nondenaturing polyacrylamide gel electrophoresis. This system is capable of separating DNA fragments that differ by as little as two base pairs. The electrophoresis unit holds two vertical 100-sample gels allowing standards and samples from a 96-well plate to be analyzed on a single gel. DNA samples are stained during electrophoresis by ethidium bromide in the running buffer. In addition, one of the gel plates is UV-transparent so that gels can be photographed immediately after electrophoresis without disassembling the gel-plate sandwich. Electrophoresis runs are generally less than two hours. The cost per gel, excluding PCR cost, is currently estimated at about $2.60, or less than $0.03 per data point. This system has been used successfully with soybean [Glycine max (L.) Merr.] and wheat (Triticum aestivum L.) microsatellite markers and could be a valuable tool for researchers employing markers in other species.


Abbreviations: bp, base pair • ITMI, International Triticeae Initiative • PCR, polymerase chain reaction • QTL, quantitative trait loci • SSR, simple sequence repeat



http://www.plantmethods.com/content/1/1/3
High throughput, high resolution selection of polymorphic microsatellite loci for multiplex analysis
Nicholas C Cryer1 , David R Butler2 and Mike J Wilkinson1

1School of Biological Sciences, University of Reading, Reading, Berkshire, RG6 6AS, UK

2Cocoa Research Unit, The University of West Indies, St. Augustine, Trinidad and Tobago

author email corresponding author email

Plant Methods 2005, 1:3doi:10.1186/1746-4811-1-3

The electronic version of this article is the complete one and can be found online at: http://www.plantmethods.com/content/1/1/3

Received: 25 May 2005
Accepted: 18 August 2005
Published: 18 August 2005

© 2005 Cryer et al; licensee BioMed Central Ltd.
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Keywords: Multiplex, Microsatellite, High Throughput, Fluorescent, Dinucleotide, High-Resolution, Allelic Ladder

Abstract
Background
Large-scale genetic profiling, mapping and genetic association studies require access to a series of well-characterised and polymorphic microsatellite markers with distinct and broad allele ranges. Selection of complementary microsatellite markers with non-overlapping allele ranges has historically proved to be a bottleneck in the development of multiplex microsatellite assays. The characterisation process for each microsatellite locus can be laborious and costly given the need for numerous, locus-specific fluorescent primers.

Results
Here, we describe a simple and inexpensive approach to select useful microsatellite markers. The system is based on the pooling of multiple unlabelled PCR amplicons and their subsequent ligation into a standard cloning vector. A second round of amplification utilising generic labelled primers targeting the vector and unlabelled locus-specific primers targeting the microsatellite flanking region yield allelic profiles that are representative of all individuals contained within the pool. Suitability of various DNA pool sizes was then tested for this purpose. DNA template pools containing between 8 and 96 individuals were assessed for the determination of allele ranges of individual microsatellite markers across a broad population. This helped resolve the balance between using pools that are large enough to allow the detection of many alleles against the risk of including too many individuals in a pool such that rare alleles are over-diluted and so do not appear in the pooled microsatellite profile. Pools of DNA from 12 individuals allowed the reliable detection of all alleles present in the pool.

Conclusion
The use of generic vector-specific fluorescent primers and unlabelled locus-specific primers provides a high resolution, rapid and inexpensive approach for the selection of highly polymorphic microsatellite loci that possess non-overlapping allele ranges for use in large-scale multiplex assays.


Frasier TR, Wilson PJ, White BN: Rapid screening of microsatellite markers for polymorphisms using SYBR® green 1 and a DNA sequencer.

BioTechniques 2004, 36:408-409. PubMed Abstract

Return to text


Narvel JM, Chu WC, Fehr WR, Cregan PB, Shoemaker RC: Development of multiplex sets of simple sequence repeat DNA markers covering the soybean genome.

Molecular Breeding 2000, 6:175-183. Publisher Full Text

Return to text


Tang S, Kishore VK, Knapp SJ: PCR-multiplexes for a genome-wide framework of simple sequence repeat marker loci in cultivated sunflower.

Theor Appl Genet 2003, 107:6-19. PubMed Abstract | Publisher Full Text

Return to text


Tommasini L, Batley J, Arnold GM, Cooke RJ, Donini P, Law JR, Lowe C, Moule C, Trick M, Edwards KJ: The development of multiplex simple sequence repeat (SSR) markers to complement distinctness, uniformity and stability testing of rape (Brassica napus L.) varieties. Theor Apl Genet 2003, 106:1091-1101. PubMed Abstract | Publisher Full Text

Morin PA, Smith DG: Non-radioactive detection of hypervariable simple sequence repeats in short polyacrylamide gels. BioTechniques 1995, 19:223-228. PubMed Abstract

Scrimshaw BJ: Non-radioactive detection of hypervariable simple sequence repeats in short polyacrylamide gels. BioTechniques 1992, 13:188. PubMed Abstract

Agarose-based system for separation of short tandem repeat loci.

BioTechniques 1997, 22:976-980. PubMed Abstract


Houriham RN, O'Sullivan GC, Morgan JG: High-resolution detection of loss of heterozygosity of dinucleotide microsatellite markers.





http://www.biotechniques.com/biotechniques/BiotechniquesJournal/2007/April/Microsatellite-marker-identification-using-genome-screening-and-restriction-ligation/biotechniques-41700.html?autnID=588199

Microsatellite marker identification using genome screening and restriction-ligation

Helena Korpelainen, Kirsi Kostamo, Viivi VirtanenUniversity of Helsinki, Helsinki, FinlandBioTechniques, Vol. 42, No. 4, April 2007, pp. 479–486 Full Text (PDF) Supplementary MaterialKorpSUPP424 (.pdf)


http://www.biotechniques.com/biotechniques/multimedia/archive/00010/97225pf01_10994a.pdf
Agarose-Based System for Separation of Short Tandem Repeat Loci


http://www.academicjournals.org/AJB/PDF/pdf2009/3Jun/Wang%20et%20al.pdf
A new electrophoresis technique to separate microsatellite alleles - QiAxcell System

http://bioinformatics.oxfordjournals.org/cgi/content/abstract/btp418v1?etoc

http://bioinformatics.oxfordjournals.org/cgi/content/abstract/btp418v1?etoc

Identification of distant family relationships

Øivind Skare 1,2, Nuala Sheehan 3 and Thore Egeland 4,5,*
1Norwegian Institute of Public Health, 0403 Oslo, Norway and 2 Department of Public Health and Primary Health Care, University of Bergen, 5018 Bergen, Norway and 3Department of Health Sciences and Department of Genetics, University of Leicester, UK and 4 Institute of Forensic Medicine, University of Oslo, 0027 Oslo, Norway and 5 Oslo University College


*To whom correspondence should be addressed. Thore Egeland, E-mail: Thore.Egeland@medisin.uio.no



Abstract



Motivation: Family relationships can be estimated from DNA marker data. Applications arise in a large number of areas including evolution and conservation research, genealogical research in human, plant and animal populations, forensic problems and genetic mapping via linkage and association analyses. Traditionally, likelihood-based approaches to relationship estimation have used unlinked genetic markers. Due to the fact that some relationships cannot be distinguished from data at unlinked markers, and given the limited number of such markers available, there are considerable constraints on the type of identification problem that can be satisfactorily addressed with such approaches. The aim of this paper is to explore the potential of linked autosomal SNP markers in this context. Throughout, we will view the problem of relationship estimation as one of pedigree identification rather than identity-by-descent, and thus focus on applications where determination of the exact relationship is important.

Results: We show that the increase in information obtained by exploiting large sets of linked markers substantially increases the number of problems that can be solved. Results are presented based on simulations as well as on real data.

Availability: The R library FEST is freely available from http://folk.uio.no/thoree/FEST.

Contact: Thore.Egeland@medisin.uio.no