Showing posts with label phylogenetics. Show all posts
Showing posts with label phylogenetics. Show all posts

Wednesday, 17 April 2013

How to read a phylogenetic tree

This week I have been preparing my last phylogenetics lectures and practical of the year. Something that is quite clear when marking student work is that many students have no idea how to read a phylogenetic tree and identify the key features about it that aid interpretation. To help combat this, here is my basic guide of "How to read a phylogenetic tree".

Topology

The first and possibly most important thing about any tree is the topology - the branching order. It's easy to get distracted by the direction or style of the tree (e.g. curved branches versus straight) but none of these things matter for the topology. The key thing here is the path taken from one node to the other and how the species (or molecules but I will just refer to species here for clarity) cluster together. The "direction" of the tree and whether the (terminal) "leaf" nodes are at the top, bottom, left or right is not important. (I'll come back to this below, for "rooting") Likewise, the vertical ordering of nodes is not inherently important - whether a particular node appears at the top, bottom or somewhere in the middle of the tree is largely a matter of preference that will depend on the purpose of the tree and the story you are using it to tell.

In fact, it is easier to look at the branches in a tree, as nodes can be rearranged in a way that can at first appear confusing. These four trees, for example, all have the same topology:
Each branch can be thought of as dividing the tree in two and splitting the species into two accordingly. If two trees share a topology, their branches will make the same splits, even if their (in this case) vertical ordering is different. Trace, for example, the path from A to D. This is easiest first in the top left tree. The branch leading directly to A splits the tree into A:BCDE. The next splits AB from CDE. (This has the root, which I will come to, below.) Now, moving back out from the root to the tip, we travel along a branch that splits ABE:CD before finally the branch leading only to D and splitting it from ABCE. Tracing the path of A to D in any of the other three trees will take exactly the same route. Any other tip to tip journeys will likewise be the same in these four trees.

This is particularly important when comparing trees, particularly big ones. I have seen people invest a lot of time and effort (and sometimes manuscript space) speculating about the differences between two particular trees when, in fact, they were really the same tree and there was no difference. Alternatively, the topology might be the same but the differences might just be due to where the tree was rooted, which I will return to.

Branch Lengths

Once the topology is clear, the next things to look at are the branch lengths, as these can give key insights into how the tree can be interpreted and, sometimes, even the methods behind the tree. There are two key things to look at in this respect: (1) the distance between (not necessarily connected) internal nodes, shown with the red arrows below, and (2) the root-to-tip distances for each terminal node, shown by the coloured arrows in the figure below:
If the spacing is even (i.e. all the red arrows are the same length) then it is highly likely that branch lengths are not being shown and the tree is only displaying the toplogy. This can be confirmed by (a) the lack of a scale bar, and (b) a bias towards internal nodes towards the tips. (In the left tree, the node joining AB is aligned with that joining CD, not the deeper CDE ancestor.) If the spacing is not even then branch lengths are being shown. These should really be accompanied by a scale bar (although the figure about does not have any).

If branch lengths are being shown, the next thing to look at is the total root-to-tip distance for each terminal node. (The coloured arrows in the figure above.) If these are all the same length, as in the right-hand tree, it is highly likely that a molecular clock has been assumed (if it's a molecular phylogeny). If it hasn't been assumed - and the methods should provide enough details to know - then the molecule in question is just evolving in an incredibly clock-like fashion. More usually, these root-to-tip distances will not all be the same. If the tree is topology-only, as in the left-hand tree, the equal root-to-tip distances do not mean anything and no conclusions about rates can be reached.

Rooting

Evolutionary trees are (almost) always starting with an ancestor and then dividing, so you can always identify the root (if there is one) as the point where all the branches converge. Historically, it was drawn at the bottom like a real tree (as with the great Molluscan tree in OUMNH and the OneZoom Tree of Life Explorer). These days, it is usually drawn on the left as in these diagrams but I have seen trees with the root at the top, bottom or even on the right. (The latter is usually only used when mirroring another tree.) I have posted before on how to root a phylogenetic tree, so I won't go over that again here. The rooting method should be given in the methods but, when it is missing, you can often guess from the shape of the tree and using the root-to-tip branch lengths again:
Unrooted trees are pretty obvious when shown in the "radiation" style. If the tree is rooted, it is almost certainly either midpoint rooted or outgroup rooted (see "how to root a phylogenetic tree"). Midpoint rooting can be identified by virtue of the fact that the two longest root-to-tip distances will (a) be the same length and (b) be either side of the root. If either of these conditions is broken, it is not midpoint rooted and is probably outgroup rooted. (Note that if both conditions are met, it is still possible that the tree is outgroup rooted. Indeed, if the evolutionary rates are fairly consistent, outgroup rooting and midpoint rooting should be the same.)

Ideally, a rooted tree should have the root marked. Sometimes, however, it is left off, as in the bottom left. This can be confusing as tree visualising programs will often display trees in the "traditional" style even when they are not rooted. This is particularly a problem when branch lengths are not shown as it will not be at all obvious when the tree is rooted or not. The time that I see this catch people out most is when making a Maximum Parsimony tree using the popular software, MEGA - these trees are displayed randomly rooted and without branch lengths by default.

Reliability and Confidence Metrics

It is always important to consider how reliable the rooting method used is likely to be if conclusions are being reached regarding the direction of particular evolutionary events. Despite this, it's pretty rare for the root position to have a direct confidence measure associated with it (although I am sure there are ways to do it). What is common, however, is to have confidence metrics for the internal branches, which are usually placed above (or sometimes below) the branch next to the descendant node (in red, below). (Branch lengths, when shown, are normally below and nearer the middle of the branch.)
Bayesian and Maximum Likelihood methods quite often produce branch probabilities as part of the method but otherwise the most common method is "bootstrapping", which is a random sampling method. I will save bootstrapping and branch tests for future posts. My one tip for now: always remember that bootstrap values are associated with branches and not nodes.

Phylogenetic checklist

In summary, my checklist for reading a phylogenetic tree: topology ⇒ branch lengths? ⇒ molecular clock? ⇒ rooting ⇒ branch confidence metrics.

Saturday, 15 December 2012

Evolution is a population-level phenomenon

An argument I have been encountering a lot recently is one that goes something along the lines of:
"Natural Selection cannot be the source of novel adaptations because it only works on what it already present. It does not generate anything and therefore a novel trait cannot be the product of Natural Selection."
These claims are then used to support the primacy of mutation as the driving force behind evolution, often coupled with (unsubstantiated) claims that these mutations are not random. In other words, Natural Selection is a myth and goal-directed mutation is responsible for the evolution of adaptations.

On face value, this can seem like a convincing argument. Natural Selection does only act on existing variation. New variants do have to arise by mutation (for a given definition of mutation), which is independent of selection. However, extrapolating that to mean that adaptive evolution occurs through the source of the new variants, mutation, and not selection is a classic case of confusing individual traits with population/species traits. I suspect that this confusion is the source of misunderstanding for many people who champion directed mutation and denigrate the power or potential of Natural Selection.

We have a tendency to draw phylogenetic trees as single lines for the branches. It is important to remember, however, that these branches - representing the evolution of species of gene sequences - are actually representing whole populations of organisms or molecules.

Evolution is a population-level phenomenon: individuals mutate but they do not evolve. It is true that new variants have to arise in an individual, independent of selection. However, we do not say that a trait has "evolved" until it reaches a high frequency or even reached (effective) fixation in the population.

For example, certain mutations cause polydactyly (extra fingers and toes) in humans but we do not say that humans have "evolved six fingers". For humans, a mutation usually has to reach a frequency of 1% before being considered a polymorphism. This is somewhat arbitrary but it needs to be in at least two generations; otherwise, lethal or sterilising mutations would constitute a polymorphism and this would make the concept pretty useless. Likewise, it would be pretty silly to say that something had "evolved sterility" because a single individual had a sterilising mutation.

Evolution does not need Natural Selection. Random processes are sufficient for a neutral (or nearly neutral) trait to "drift" its way through a population to fixation. However, without invoking an external agent, Natural Selection is the only process that drives a trait through (or from) a population, resulting in adaptive evolution.

Population-level change is still change. A population or species acquiring a new trait is still evolution of a new trait. A novel evolved trait can be the product of Selection.

Monday, 3 December 2012

Zooming around the Tree of Life


This morning, whilst feeding the cats, I came across the very fun - and educational - OneZoom Tree of Life Explorer, described in a PLoS Blog article, Fractaltastic Evolution. Currently, it only has mammals and amphibians but it will grow. Birds are next and plans are afoot to use Open Tree of Life data to extend it to 2 million species (or maybe more by then).

It has lots of nice features, including different views, threats of extinction from the IUCN Red List and dates of divergence. The latter can be used to run a "Growth Animation" timeline, which is another useful tool for trying to grasp evolutionary timescales. I can feel some OneZoom-inspired MapTime TimeLines coming on when time allows.

h/t: @phylogenomics

Friday, 27 July 2012

Marvellous Mollusca

Arestorides argusIt's not just rocks that have crazy and beautiful patterns - shells do too. The Oxford University Museum of Natural History has a great collection of such shells and this one stands out as a particularly fine example. It's labelled Cypraea argus, now known as
Arestorides argus - a.k.a. the "eyed cowrie", a species of sea snail.

Whilst undeniably pretty, the eyed cowry is not my favourite part of the Mollusca exhibit at OUMNH, though. That honour goes to the phylogenetic tree of Molluscan Classes made out of mollusc shells. How cool is that‽ (As a marine gastropod, Arestorides argus would join this particular party in the top left.)
mollusca treePerhaps the craziest molluscs of all, however, are the Cephalopods, which include octopuses and squid (including both flying and glow-in-the-dark varieties). These guys have lost their shells (or only retain small internal parts), though, so they can't contribute to this particular tree.

Thursday, 7 June 2012

How to root a phylogenetic tree

Teaching phylogenetics, it is clear that one of the things that causes a surprising amount of confusion is rooting the tree - defining the position on the tree of the (hypothetical) ancestor. Here, then, is my basic introduction to rooting a phylogenetic tree and why it is important.

Unless you are modelling recombination or Horizontal Transfer, a phylogenetics is explicitly based on a model of a single common ancestral lineage that splits over time into multiple lineages. For species, these splits are speciation events. For molecules (e.g. genes or protein sequences), these splits can also be duplication events. But how do you know where that common ancestral lineage is?

With respect to rooting a phylogenetic tree, there are three main strategies that are routinely employed. In approximate order of confidence in the ancestry, these are: (1) No rooting (leave the tree unrooted), (2) Midpoint rooting, and (3) Outgroup (not outlier!) rooting.

Unrooted. I'll start with the unrooted tree because most methods produce an unrooted tree. Note that, although phylogenetics as a general approach assumes there's a common ancestor somewhere, many of the actual methods for constructing trees from data make no direction/ancestry assumptions. (The obvious exception is UPGMA, which is not used that much for serious phylogenetics.)

Leaving the tree unrooted is also the easiest because it essentially involves doing nothing! In the figures below, the left-hand image shows the unrooted tree. The key thing here is that there is no assumption about ancestry and therefore no statement about the direction of evolution. If you trust the approximate branch lengths, it is still possible to say things about the topology, such as A and B are closer to each other than they are to C and D (with an important caveat covered below), but you cannot make any direct inferences about the common ancestor. Whilst easier, therefore, you are limited in terms of analysis and interpretation.

This kind of representation is best used in situations where you have three or more well-defined clades (such as different members of a gene family) that radiated a long time ago over a relatively short period, such that the placement of the root and the order of branching is not known. You might still be able to say something with confidence about some of the individual radiations but you avoid making unsupported claims about the precise history and relationships of those clades. (Another option here is to "collapse" the root into a multifurcation (one-to-many split) but I will not deal with that here.)

The alternative when the actual root placement is unknown, is to use midpoint rooting:



Midpoint Rooting. As its name suggests, Midpoint rooting attempts to root the tree in its middle point. It does this by calculating all of the tip-to-tip distances and selecting the longest - A to E in the tree above. The root is then placed exactly half-way between these two tips. If the tree is behaving and the rates of evolution are pretty constant throughout, this point should represent the ancestral point. It is therefore useful in situations where the actual root is not known but the assumption of a reasonably constant "clock-like" rate of evolution is quite sound. The other situation it generally works well for is when a tree is fairly balanced with some closely-related clades separated by a long branch in the middle - if midpoint rooting places the root far away from any nodes, it is less likely to be wrong (i.e. moving it a little due to rate discrepancies would not make any difference).

The main problem with midpoint rooting is that it is very susceptible to large deviations from a constant evolutionary rate, especially if these are not "balanced", i.e. they only occur on one side of the actual root. The other time midpoint rooting tends to go wrong is when it places the root in amongst a rather dense set of short branches, where quite small deviations will place the root on another branch. For these reasons, whenever possible, outgroup rooting is generally the method of choice.



Outgroup Rooting. Unlike midpoint rooting, in which features of the tree itself specify the root, in outgroup rooting existing knowledge is utilised to place the root in the right place. This is done by using an "Outgroup" - a species or molecule that is known to be more distantly related than everything else in the tree. In the example above, kangaroo is used to outgroup root the tree, as it is known that marsupial mammals diverged from the ancestral lineage of all placental mammals.

This tree emphasises the point made above about evolutionary rates - the midpoint root was wrong because the rodent lineages (mouse, rat and their ancestor) are evolving faster than the rest of the tree, probably due to relatively short generation times. (This pattern is often seen with real trees.) This can sometimes be obvious if deleting one of the nodes used to midpoint root the tree changes the position of the root: for a perfect clock-like tree, it should make no difference (as long as you do not delete the outgroup). This will not always work, though - in the example above, it would not make a difference, for example.

Why does the root matter? There are a couple of reasons why correct rooting is important. The first just comes down to interpretation and understanding - it would be wrong to get too obsessive about the superficial differences between the three trees in the above figure - they are all essentially the same tree (same topology) and the differences are all down to rooting. The second is more important and comes down to the direction of evolution and conclusions about relatedness. If you want to infer anything about ancestry, you obviously need to have the right root. From the midpoint rooted mammalian tree above, it might not be obvious that placental mammals form a Monophyletic clade. Coming back to my earlier caveat when interpreting an unrooted tree, you might determine that the kangaroo, lemur and human were all more closely related to each other than to the mouse and rat. In terms of pure sequence divergence (on which the tree was built) this might be right but, in evolutionary terms, it is wrong: all placental mammals are equally distant from kangaroos due to the shared common ancestor.

Related post: How to read a phylogenetic tree.

Tuesday, 5 June 2012

A Molecular Evolution Glossary

After my earlier Outgroup/outlier, I have made a brief Molecular Evolution Glossary that I made for some of my lectures available. It was made predominantly so that the students knew what I was meant by various terms in the lectures - Outgroup is in there, so I'm not sure how much they used it! - but it might be of use to someone beyond my students. If so, feel free to use it - hopefully it can save someone else the effort of putting together something similar. Please let me know if you spot anything obviously wrong or confusing (or missing) and I'll try to update it from time to time. (I'm sure I have a few terms to add from other glossaries that I've assembled over the years.)

Wednesday, 30 May 2012

Phylogenetics: Outgroups and Outliers

I am in the throes of marking at present and one mysterious error is cropping up enough to warrant a short blog post: the confusion of the terms "outgroup" and "outlier" when discussing phylogenetic trees. The source seems to be a familiarity (or exposure) to the term "outlier" in statistics, combined with a lack of familarity/awareness of the term "outgroup". So, what is an outgroup?

An outgroup in phylogenetics simply refers to a lineage that is known to be more distantly related to the other species (or DNA/proteins) being studied. In the example shown, the kangaroo is a known outgroup to the other mammals, as marsupials diverged from the common ancestor of placental mammals prior to the subsequent radiation of placental mammals. This knowledge can be used to "root" the tree, as in this example, but more on that in a later post.

What about outliers? "Phylogenetic outliers" are sometimes referred to but, as far as I am aware, there is no standard definition. Instead, it is a term that is used when some form of numerical test (such as evolutionary rates) or statistics have been applied to a phylogenetic tree and there are particular branches that do not fit the norm - they are "outliers". In this tree, the kangaroo is not noticeably an outlier, although it's evolutionary rate depends on a somewhat arbitrary position of the root on the kangaroo branch. The rodents might be outliers (if there were more data) as they tend to evolve quite quickly due to short generation times. In this particular tree, I doubt it though. (This tree is made up, by the way.)

Most important, though, "outlier" is not a synonym for "outgroup" and should be used with great caution.

Wednesday, 9 May 2012

Education is the key to impact

This month is a busy one and two activities that were going on side-by-side last week were (1) writing an outline grant proposal to further develop SLiMSuite (servers and programs) and SLiMdb, and (2) dig up possible "impact" for the upcoming REF. Part of this has featured digging into Google Analytics data for both the bioware servers and my homepage.

Sadly, I stupidly neglected to monitor the documentation and download pages for SLiMSuite (now rectified) but I could get visit data from the webservers and the results made fairly happy reading, especially when compared to the citation metrics associated with the relevant papers.

The bioware.ucd.ie front page currently receives around 350 visits a month from across the world, with 4,174 visits in the period 1 May '11 to 30 Apr '12. Ireland is currently the main user, reflecting the fact that the current host of these resources is University College Dublin, although the balance seems to be shifting towards America now. (Given the amount of science being done there, you would expect the US to be number 1.) Germany and the UK complete the top 4 (again not surprising in terms of both science and presence of participating/collaborating labs) but the usual suspects (India, Canada, Israel, China) are all in there too, which is good to see. The servers themselves receive around 500-1000 page views a month, with the top five servers visited in the period 1/5/11-30/4/12 being: 1. SLiMFinder (2,651 views); 2. SLiMPred (1,812 views); 3. SLiMSearch 2.0 (1,721 views); 4. SLiMSearch 1.0 (805 views); 5. CompariMotif (653 views).

This is all well and good - these are published servers - and hopefully will continue to increase over time, particularly if we can get more funding for development. The thing that surprised me, though, was when I looked at the usage stats for the UPGMA walkthrough I made for a Year 2 practical that I run with a colleague. This website got 5,721 visits in the same period (1 May '11 to 30 Apr '12) - 37% more hits than the bioware front page. This number is also on the rise with nearly 1000 hits last month, due in part to the fact that it now appears on the UPGMA wikipedia page.



This perhaps should not be surprising but when you consider that UPGMA is a largely obsolete method used predominantly for teaching (and quick and dirty clustering) and I knocked up the website in an evening or two, while SLiMFinder is a cutting-edge bioinformatics tool that represents years of hard work, it is still a little depressing. I guess I should make more educational web pages...