Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Friday, 27 January 2017

In Defense of Science

From the Evolution Directory (evoldir) mailing list today:

Governmental scientists employed at a subset of agencies have been forbidden from presenting their findings to the public. We have drafted the following response for distribution, and encourage other scientists to post it to their websites, when feasible.

Graham Coop
Professor of Evolution and Ecology
UC Davis

Michael B. Eisen
Professor of Molecular and Cell Biology
UC Berkeley

Molly Przeworski
Professor of Biological Sciences
Columbia University


The message, for any affected US scientists out there:

We are deeply concerned by the Trump administration’s move to gag scientists working at various governmental agencies. The US government employs scientists working on medicine, public health, agriculture, energy, space, clean water and air, weather, the climate and many other important areas. Their job is to produce data to inform decisions by policymakers, businesses and individuals. We are all best served by allowing these scientists to discuss their findings openly and without the intrusion of politics. Any attack on their ability to do so is an attack on our ability to make informed decisions as individuals, as communities and as a nation.

If you are a government scientist who is blocked from discussing their work, we will share it on your behalf, publicly or with the appropriate recipients. You can email us at USScienceFacts@gmail.com.

I’ve also heard a rumour that Michael Eisen is running for senate. That would be cool - we need fewer Trumps and more science-savvy politicians.

Thursday, 26 January 2017

Is it fair to compare "Alternative Facts" with "Alternative Medicine"

I came across this “meme” on Facebook today:

“The very concept of alternative medicine exists to create a double standard where the rules of science and evidence are stood on their head specifically to manufacture the result that is desired by cranks, charlatans, snake-oil salesmen, and self-proclaimed gurus. There is no alternative medicine. There is just medicine. Either it works or it doesn’t work.”
-Steven Novella

The inevitable retort was: but what about some “natural/traditional” medicines that have not been thoroughly studied. These might work. So, surely it’s an unfair comparison?

My response to that is this:

When people bullshit based on their gut feelings, they can also be right sometimes. It’s only when someone looks into it do we know whether it is actually fact or fiction. “Folk medicine” and natural products may have good (or bad!) effects but it is wrong to imply that they are medicine until we know whether/when they work.

In the same way that an opinion is not an “alternative fact”, a natural product that somebody thinks might do something is not “alternative medicine”.

Then, of course, there is the less generous - but even more apt - comparison of bare-faced lies with bare-faced fraudulent treatments like homeopathy - things demonstrably false that are being badged at truth under the label “alternative”.

I just hope that the war on “Alternative Facts” is more successful than the war on “Alternative Medicine”. The real problem with taking action based on made up stuff is that reality doesn’t care how well-meaning you are, or how much you want it to be true. Hopefully, America will not suffer too much at the hands of reality before Trump and/or his cronies realise this.

Sunday, 27 March 2016

Are data scientists just "research parasites"?

Although it passed me by at the time, the New England Journal of Medicine - a highly respected top-tier medical journal - featured an editorial on data sharing1 in January. It was so bad, that the International Society for Computational Biology (ISCB) felt the need to respond in the most recent issue of PLoS Computational Biology2. I’m glad they did, for the editorial was awful.

It starts quite well:

The aerial view of the concept of data sharing is beautiful. What could be better than having high-quality information carefully reexamined for the possibility that new nuggets of useful data are lying there, previously unseen? The potential for leveraging existing results for even more benefit pays appropriate increased tribute to the patients who put themselves at risk to generate the data. The moral imperative to honor their collective sacrifice is the trump card that takes this trick.

But then rapidly goes downhill:

However, many of us who have actually conducted clinical research, managed clinical studies and data collection and analysis, and curated data sets have concerns about the details. The first concern is that someone not involved in the generation and collection of the data may not understand the choices made in defining the parameters. Special problems arise if data are to be combined from independent studies and considered comparable. How heterogeneous were the study populations? Were the eligibility criteria the same? Can it be assumed that the differences in study populations, data collection and analysis, and treatments, both protocol-specified and unspecified, can be ignored?

Many of us who have actually conducted data analysis would retort: if you have concerns about the details then you should be making those details clear. If choices are important, explain them! For sure, you cannot just blindly combine multiple datasets that have different biases etc. but what decent scientist would do that (without an explicit caveat regarding that assumption)?

Longo and Drazen seem to be implying that all data scientists are bad scientists. As I’ve said before, Bioinformatics is just like bench science and should be treated as such. If you are making dodgy assumptions about data, you are doing it wrong. (Though people do make mistakes - the data collectors too.)

It gets worse:

A second concern held by some is that a new class of research person will emerge — people who had nothing to do with the design and execution of the study but use another group’s data for their own ends, possibly stealing from the research productivity planned by the data gatherers, or even use the data to try to disprove what the original investigators had posited. There is concern among some front-line researchers that the system will be taken over by what some researchers have characterized as “research parasites.”

Apparently, some people might think I am a “research parasite” because I sometimes analyse other people’s (published) data without talking to them about it. I’m glad the ISCB called them out on this. Newsflash: science only makes progress by people trying to disprove what other researchers (and, ideally, themselves) have posited. Science is a shared endeavour. If someone uses your data to do something (good), good! If you don’t want that, embargo the data or delay publication. Then question your motives; if glory is what you seek, perhaps you’re in the wrong profession?

A researcher frightened of “stolen productivity” is perhaps a researcher struggling for ideas. (I’d love someone else to answer some of the questions I have kicking around so that I could move on to the next thing!) A researcher scared of someone trying “to disprove what the original investigators had posited” has bigger problems.

The rest of the editorial is not so bad, as it tells the tale of a fruitful collaboration between “new investigators” and “the investigators holding the data”. Of course, this is the ideal scenario, short of generating the data themselves. The fact that the authors felt the need to stress this - and the language used of “symbiosis” versus “parasitism” - demonstrates that Longo and Drazen are utterly clueless about the modus operandi of the disciplines they discredit. Whilst ideal, direct collaboration is not always feasible. Sometimes - when the original investigators are too attached to their pet hypothesis or conclusion - it is not desirable.

They end:

How would data sharing work best? We think it should happen symbiotically, not parasitically. Start with a novel idea, one that is not an obvious extension of the reported work. Second, identify potential collaborators whose collected data may be useful in assessing the hypothesis and propose a collaboration. Third, work together to test the new hypothesis. Fourth, report the new findings with relevant co-authorship to acknowledge both the group that proposed the new idea and the investigative group that accrued the data that allowed it to be tested. What is learned may be beautiful even when seen from close up.

This sounds OK - and the described model may even be data sharing at its best - but the implication that anything short of this ideal is somehow inadequate is naive and unhelpful.

First, one person’s novel idea is another person’s obvious extension. And anyway, why should having one idea give you automatic rights to all obvious extensions?! Why should the rest of us trust the data gatherers to do a good job - especially if they exhibit attitudes towards data akin to these authors?

Second, identifying a potential collaborator does not guarantee collaboration. Ironically, the kind of paranoid narcissist that would use a term like “research parasite” is unlikely to be open to collaboration.

Thirdly, citation is a form of co-authorship that acknowledges “the investigative group that accrued the data”. Wanting full co-authorship where additional intellectual input is not required is just greedy. (And a note to the narcissist: self-citations are generally seen as lower impact than citations by wholly independent groups.)

Longo and Drazen should stick to commenting on what they know, whatever that is, and leave data scientists to worry about how they conduct themselves. With this editorial, they have done everyone - not least of which themselves - a deep disservice.


  1. Longo D.L., Drazen J.M. Data Sharing. N Engl J Med, 2016. 374(3): p. 276–7. doi:10.1056/NEJMe1516564.

  2. Berger B, Gaasterland T, Lengauer T, Orengo C, Gaeta B, Markel S, et al. (2016) ISCB’s Initial Reaction to The New England Journal of Medicine Editorial on Data Sharing. PLoS Comput Biol 12(3): e1004816. doi:10.1371/journal.pcbi.1004816.

Saturday, 26 March 2016

Meet the world's newest lifeform: Syn 3.0

Every now and then, a piece of science is done that is truly ground-breaking and world changing. One such piece is:

Hutchison III CA et al. (2016) Design and synthesis of a minimal bacterial genome. Science 351(6280): aad6253-1. DOI: 10.1126/science.aad6253

Science has a summary here but it’s worth reading the whole paper. Syn 3.0 itself is pretty impressive, but what’s even more impressive is the approach taken to make it. In addition to using current knowledge of fundamental biological machinery, the Venter group used large-scale transposon mutagenesis and selection to identify additional genes that were either essential (i.e. no growth without them) or “quasi-essential”, where removal resulted in a major growth deficit.

They also had to overcome the problem of redundancy: even in a genome as reduced as the Mycoplasma species, there can sometimes be multiple genes that do the same thing. Removing one makes little difference but removing both is lethal - something hard to identify when knocking out single genes at a time. Whatever the Intelligent Design crowd would like to believe, biology is messy.

Of course, Syn 3.0 is just the start, as the goal was making a “minimal cell”:

“A minimal cell is usually defined as a cell in which all genes are essential. This definition is incomplete, because the genetic requirements for survival, and therefore the minimal genome size, depend on the environment in which the cell is grown. The work described here has been conducted in medium that supplies virtually all the small molecules required for life. A minimal genome determined under such permissive conditions should reveal a core set of environment-independent functions that are necessary and sufficient for life. Under less permissive conditions, we expect that additional genes will be required.”

Robust life will therefore need a lot more genes. It will be interesting to see how many are required for autotrophy - life that needs only inorganic chemicals and an energy source.

Even within the “minimal cell” concept, Syn 3.0 represents a somewhat arbitrary end-point. In identifying the “quasi-essential” genes, a judgement had to be made regarding what constitutes an acceptable growth rate*. Whittling down to 473 genes is impressive, but this number could no doubt be even smaller if slower growth rates were accepted. (Modern life is in competition with lots of other highly evolved organisms. Early life would have been able to get by with much lower growth rates, so this is not a “minimal cell” in that context.)

There is also a lot of exciting potential ahead for manually reducing the number of genes by true intelligent design. Fusing interacting gene products together, for example, might eliminate the need for so many genes contributing to core processes. (Looking for apparent protein fusion/fission events in evolution is a reasonably successful method for predicting protein-protein interactions.) With time, we might be able to “wind back the clock” and remove some of the unnecessary complexity that has probably crept into the system due to the underlying evolutionary process.

I also wonder how many of the current crop of genes of unknown function - a surprising 149 genes - can be replaced over time with genes of known function. (In other words, how many of them represent convergent evolution of functions we already know about but are not recognisable.) And how many of the rest are genome-/condition-specific?

Like all of the best science, this work opens the door to more questions than it answers! Some exciting times ahead, I think.

[*The important but oft-overlooked concept that any assessment of life is context- and environment-dependent exposes another flaw with Intelligent Design as a testable hypothesis: designed to do what? To assess how well-designed something is, one needs to know its purpose and/or the acceptable design traits. To hide from the fact that Intelligent Design is Creationism, supporters often make the argument that the identity of the designer (Creator) is not important - but without knowledge of the designer, how can one predict the motivation behind the design?]

Monday, 29 February 2016

Thank YOU, PLOS ONE!

There are many flaws with the peer review system but it remains the best system we have for ensuring a certain degree of quality control prior to publication. One of the ways that the system could be improved is better recognition - and therefore motivation - for reviewers.

Ideally, there would be some form of payment, but I find it hard to see this happening any time soon. (It is difficult enough to get funds to publish papers - getting funds to get papers reviewed when they might well end up getting rejected is going to be way harder.)

The next best thing is some kind of reward or recognition. Some journals give discounted publication fees to reviewers, which is a great idea. Another great idea has just been put into action by PLOS ONE*: public recognition for reviewers:

On behalf of PLOS and the PLOS ONE editorial team, I would like to thank you for participating in the peer review process this past year at PLOS ONE.

We know there are many claims on your time and expertise and we very much appreciate your valuable input in 2015. With your help, we have continued to publish an influential, lively and highly accessed Open Access journal. Simply put, we could not do it without you and the thousands of other volunteers for PLOS ONE and the other PLOS journals who graciously contributed time reviewing manuscripts.

A public “Thank You” to our 2015 reviewers – including you – was published earlier this week.

(2016) PLOS ONE 2015 Reviewer Thank You. PLoS ONE 11(2): e0150341. doi:10.1371/journal.pone.0150341

Your name is listed in the Supporting Information file associated with the article. I hope that you will be able to use this letter, along with the article citation, to claim the credit and recognition you deserve within your institution for supporting PLOS ONE and Open Access publishing.

The article itself is short but sweet:

PLOS and the PLOS ONE editorial team would like to express our gratitude to all those individuals who participated in the peer review process of submissions to PLOS ONE over this past year. During 2015 PLOS ONE published over 28,000 research articles. This would not have been possible without the contribution of more than 76,000 reviewers from around the world and a wide range of disciplines.

The names of our 2015 PLOS ONE reviewers are listed in S1–S5 Reviewer Lists. Thank you to all our reviewers for generously sharing your time, insight and expertise with PLOS ONE authors in the evaluation of their work. Your efforts are a key reason for PLOS ONE’s success as an innovative and influential publication.

It’s nice to be appreciated. One more reason to be a fan of PLOS ONE. (Which I am, despite those who look down their noses at the journal because of its “scientifically rigorous research, regardless of novelty” policy.)

*The other PLOS journals did it to but I did not review anything for them this year.

Monday, 2 November 2015

ICBCSB 2015: Another scam conference comes to Australia

I am a bioinformatician working in Sydney, Australia. I recently helped to organise the ABACBS2015 conference in Sydney, where ABACBS stands for the Australian Bioinformatics And Computational Biology Society, of which I am a member. I am also part of the New South Wales Systems Biology Initiative. You may therefore find it surprising to know that I found it surprising to find out that in December, Sydney NSW will be host to "ICBCSB 2015 : 17th International Conference on Bioinformatics, Computational and Systems Biology".

This is not the only surprising thing about ICBCSB 2015. For Sydney is actually at least the 15th 17th International Conference on Bioinformatics, Computational and Systems Biology (ICBCSB 2015). Next week, the conference is being held in Madrid. Last month, there was the 17th International Conference on Bioinformatics, Computational and Systems Biology in Bali. And Prague. And Chicago. And Istanbul. The 17th International Conference on Bioinformatics, Computational and Systems Biology (ICBCSB 2015) has also been in Venice, London, New York, Berlin, and Geneva and will be held in Penang and Dubai. And that’s just in the first two pages of a Google Search - there are more (including Lisbon and Stockholm).

This makes OMIC Group Conferences look positively legit. Indeed, the organisers of all 15+ ICBCSB 2015 conferences - the World Academy of Science, Engineering and Technology (WASET) - seem to be basing their conference business model on that of OMICS. Or perhaps it was the other way round. Either way, if you replaced the WASET logo on the website with OMICS Group, nothing would seem out of place.

If you have ever attended a (real) scientific conference, pick one of those past conferences at random, for example Berlin, and click on the conference photos page, then tell me if you have ever seen anything so depressing in your life. And remember: these are the photos they chose to put up, so presumably show the conference in its best light. Given the number of group photos of (all?) the delegates, I wonder whether they had time for much else other than publicity shots. And in case you are thinking: “those are probably just the invited speakers”, I would bet good money that the delegates were all invited - and still had to pay top dollar to attend.

Finally, in case you have any doubt, just Google “WASET scam”. It does not make for happy reading.

If you are in bioinformatics or systems biology, please spread the word far and wide about these conferences and why they should be avoided at all costs. The sooner we starve the likes of OMICS Group and WASET of naïve unsuspecting scientists to prey on, the sooner these parasites will f#@k right off. There are plenty enough legit conferences to choose from. (Yes, it makes me angry.)

And if you have already signed up for Sydney ICBCSB 2015, do not despair. As luck would have it, there is a real bioinformatics event being held in Sydney that same week: BioInfoSummer 2015. It’s more of a workshop than a conference but there will be plenty of opportunities to discuss science with some excellent bioinformaticians and systems biologists. At least that way, you won’t waste the plane ticket and hotel costs, even if WASET won’t give you a refund for pulling out. (Not likely!)

Sunday, 1 November 2015

A tale of two conferences - #ABACBS2015 and #AGTA15

Two weeks ago, I attended the awesome double-header of ABACBS2015 and AGTA15. In some ways they were chalk and cheese - ABACBS was cheap and cheerful, where as AGTA was expensive and (therefore) exclusive - but both were great and highlighted some of the very best features of successful academic conferences.

As part of the ABACBS2015 organising committee, I needed to collect my thoughts for a debrief, so I thought I’d got down some thoughts here. (Thanks to grant writing, the post itself got a little delayed!) As with biology itself, most decisions are trade offs and the following comparison is not to criticise - I just find it interesting. As with the awesome ABiC14 conference last year, I mainly want to document good practice for future reference. (Some of these will no doubt be repeats of my ABiC14 thoughts!)

Size. Both conferences were, for me, the perfect size - around 180-190 people. This is enough to give the conference a good buzz and ensure that there are sufficient interesting people to listen to and posters to visit. Critically, though, it was small enough that you could find and speak to the people you wanted to. (Even if I didn’t fully get the chance at ABACBS2015 because I was busy, and managed to miss some people at AGTA15 because time ran out!)

Venue. ABACBS2015 was run on a budget and we were very fortunate to have the Garvan Institute provide a free venue for the conference. The Garvan lecture theatre is lovely, with a great AV system and comfy seats. The only negative for a conference of that size - and we got close to our 200 person limit - is that the space outside the theatre for breaks and posters does get a bit cramped. In contrast, AGTA15 was held in the Crowne Plaza at Hunter Valley, which is a pretty luxury hotel in the wine region of New South Wales. It took me a while to warm to the main conference setup, with seats around tables in a large, wide room with a screen each side of (and a long way from) the podium. However, the “trade exhibition” space where the breaks, lunch and posters were, was excellent. I am big fan of having the posters up all the time and in the same place as the breaks.

Location. ABACBS was held in the city centre at a research institute. This kept the costs down and made it easy to get to. (We did not organise accommodation.) It also made it easy to escape from, which might have affected evening social numbers. AGTA was out in the sticks, which made it harder to get to and may have put some people off attending - it essentially added another day to the length of the conference. The plus side of this was that everything was on one site and it was not so easy to disappear or go home, and so most people were around most of the time, I think. (Although there was the temptation of wine tasting on the doorstep!)

Goodie bags. I’m pleased to say that both conferences went down the reusable shopping bag route (pick above). Being budget, we got UNSW to sponsor us with some canvas bags for ABACBS. (Complete with the scary statistic that plastic bags take 15-1000 years to break down.) AGTA was (a) a bit more upmarket and (b) in wine country, so they provided wine coolbags!

Trade exhibition. We binfies are a cheap lot - and I don’t mean tacky or miserly, I mean that we don’t need a lot of money to do great things. As a result, it is hard to attract trade sponsors: we just don’t spend (or often even have!) any money! Genomics is clearly a different matter, as you literally cannot do it without lots of expensive kit. The AGTA trade exhibit was therefore pretty big and bustling. Again, putting it with the food and posters made for some great…

…Breaks. Breaks really do make or break a conference. One of the few criticisms of ABACBS2015 was that the breaks were a bit too short - we were somewhat hemmed in for time because we needed to leave people enough time to get to AGTA in the afternoon of the second day. We also tended to over-run a little, because people would be enjoying the breaks and take a bit of time to filter back into the theatre for the talks. My advice for conference planners: (1) make all breaks 5-10 minutes longer than you think they should be (e.g. 40 minutes for coffee and over an hour for lunch); (2) build in some dummy time into the program to soak up delays; (3) Make sure you have a bell or something to signal that the next session is starting! Being a more leisurely multi-day conference, AGTA had nice long breaks, including a session off for posters etc. after lunch. Lots of opportunities for mingling and looking at posters.

Invited speakers. Both conferences had outstanding invited speakers that gave really interesting talks. I know some of the binfie crowd would have enjoyed a bit more about the methodology in ABACBS2015 - and we should perhaps brief our speakers a little better in future - but personally I was blown away by the quality of the science presented. Both fields are rapidly changing as technology opens up opportunities and there is some really cool stuff going on out there! AGTA in particular seemed to have a lot of invited speakers, talking about really cool stuff - possibly part of the reason it's so expensive. (It must be said: the quality of the selected abstracts was good too!)

Gender Equality. Both conferences did pretty well on the gender front, I think. ABACBS2015 in particular nailed it, with a 50% split of invited speakers and over 50% female speakers overall! (The latter wasn’t deliberate as such - there are just loads of good female bioinformaticians in Australia who submitted interesting abstracts!)

Twitter. (And #confBingo.) I have mixed feeling about live tweeting at conferences. This year, I decided to throw myself in a bit more and, on balance, I feel that I got more out of it than I missed. It is true that sometimes I missed something a speaker was saying because of a tangential Twitter conversation (such as #JediKelpie) - but I also learnt stuff and picked up on things that I had missed thanks to the Twitter feed. The #confBingo thread was also quite entertaining a fun, and helped networking and building community spirit, which is what a lot of a good conference is ultimately about.

Coffee. Unfortunately, we were unable to attract the barista sponsor from ABiC14 to sponsor ABACBS2015. AGTA did have a sponsored barista, though. Roche (I think) gave out three tickets with registration for the trade exhibit barista. Not quite as good as unlimited but it hit the spot. My other top conference coffee tip: bring a reusable cup! It' so much easier/safer if you want to take coffee into the talks and generally cuts down on spills - and refills! (Next year, I am hoping for ABACBS Keep Cups in the goodie bags!)

Overall, I had a lot of fun helping with the conference organisation - and would recommend it - and even more fun attending. Australia has some great science going on and a really strong/vibrant bioinformatics community, and I am proud to be part of it.

Friday, 21 August 2015

Bioinformatics is just like bench science and should be treated as such

A bad workman blames his tools. A bad life scientist blames bioinformatics. OK, so that’s a little unfair but so is the level of criticism levelled at bioinformatics by people who should know better. If you are a bioinformatician, it is inevitable that you will run up against the question of whether you ever do “real” science.

If you are unlucky, it will be as blunt as that. At the end of a bioinformatics seminar earlier this year, someone actually asked (in what was meant to be a good-natured way): “Is bioinformatics real? I give the same data to two different bioinformaticians and get completely different answers!” Often, it is is in the subtle form of: “are you going to validate that in the lab?” - as if validating it another way would not itself be valid.

If you are a bioinformatician and are asked a question like that, the correct answer is: it’s as real as [insert appropriate “wet” discipline of choice]. For bioinformatics is science and like all science it can be done well, or it can be done badly. It can generate meaningful results, or meaningless results.

If you want to get meaningful results, you have to treat it like a science, rather than a “black box” of magic. What do I mean by that? Here are my not-quite-buzzfeed-worthy, “8 shocking ways that bioinformatics is just like bench science”:

1. Experience and Training. You wouldn’t hand someone a wet lab protocol for, say, Southern blots, give them a key to your lab and say, “off you go” without first giving them some training and, preferably, the opportunity learn from someone who already knows the procedure. If you do, expect bad results. Bioinformatics is no different. Just because a two year old can work a computer these days, that does not mean that bioinformatics is easy. A two year can also press “Start” on a PCR machine.

2. Optimising your workflow. If you were doing a PCR, you wouldn’t just find a random paper you like that also did a PCR, copy the Mg2+ concentration etc. and then bung it in PCR machine and run the default cycle. Likewise, you should not just stick your data into a bioinformatics program and automatically expect it to do the right thing. Just as to be a good molecular biologist, you need to be (or know) someone who knows a bit of chemistry to understand what’s going on, to be a good bioinformatician, you need to be (or know) someone who knows a bit of molecular biology (and chemistry! and sometimes physics) to understand (a) the data you are putting into a program/workflow, and (b) what the best way to process that data is. If you make the wrong assumptions of your data, you will get the wrong answer. (And if different people make different assumptions, they will probably get different answers.) Computers just do what they are told - don’t blame them if you tell them to do the wrong thing. (It is also important not to get “target fixated” on perfect optimisation; just like for bench science, the performance of your bioinformatics workflow only needs to be as good as your experiment/question demands.)

3. Planning. You wouldn’t start a bench experiment without planning it first. Just because bioinformatics is not time-dependent, that doesn’t mean that you shouldn’t plan your analysis before you start. Know what your final output is going to be and work backwards. Making decisions as you go along is a great way to make bad decisions. Sure, have a play to work out how things work but then go back and do it properly from beginning to end.

4. Lab notes. You wouldn’t just stick a tube labelled “20/8/15 mouse 3 PCR” in the freezer and expect to remember what it was and how it was made 3 months later. Instead, you would (hopefully!) keep a meticulous record of primers and reaction conditions etc. in your lab book. Bioinformatics needs the same record-keeping mentality. Program version numbers, dates and settings are important. Write them down. You will almost certainly end up running an analysis more than once, and it won’t always be the last run that you end up using. You do not want publications to be held up because you are having to re-run your bioinformatics just to work out what settings you settled on.

5. Labelling. Even with a well kept lab book, you wouldn’t store samples or extracted DNA in tubes labelled “tube 1”, “tube 2” etc. for every experiment. If someone rearranges your freezer - or there is an emergency freezer swap following power/equipment failure - you could quickly get muddled up. Likewise, don’t call your files things like “sequence.fasta”. You’re just asking to accidentally analyse the wrong data. Include multiple failsafes so that if you enter the wrong directory, for example, your file names won’t be found. (A pet hate of mine is bioinformatics software that outputs the same generic file names each time it is run for lazy scripting.)

6. Reproducibility. Bioinformatics is - or should be - extremely reproducible in a trivial way. If you put the same data into the same program with the same settings, you should get the same answer. In this sense, it should have the edge over bench work - you do not need to repeat experiments. Right? Well, not really. Just because it should be consistent in how it goes wrong, you cannot be sure that a bioinformatics tool is not getting confused by some subtle nuance or peculiarity of your data. Try with another tool that does the same job, or change a setting that should make no difference, and check that you get qualitatively the same answer. (Better still, try changing a setting that should make a predictable difference and make sure that it does.) Just as you can get a misleading lab result if you mislabel your tubes or add the wrong buffer, you can get a misleading bioinformatics result if you mislabel your data or use the wrong parameter settings.

7. Validation. Bioinformatics often receives a certain amount of flak that it is not “real” and everything needs to be validated. This is true, up to a point. The forgotten point is that nothing is real and everything needs to be validated. In the lab, you rarely actually measure or observe something directly - you are inferring reality from things you can measure (e.g. fluorescence) based on what you think you know about the system (e.g. what you’ve labelled) and certain assumptions (e.g. lack of off-target binding). You then have to perform additional experiments - and/or bioinformatic analysis of your data - to test that your assumptions appear to be good and that there are not alternative explanations for your observations. Bioinformatics is no different. NO different. You make assumptions and you make inferences based on observed outputs. These assumptions and inferences need to be tested. This might be by “validation” in the lab. It might be by independent analysis of other data. The only reason the former is more common is that one often needs to generate new data, which clearly bioinformatics cannot do. However, if the data already exists, there is no reason why bioinformatics cannot be used to validate other bioinformatics, or even bench experiments.

8. Limits. Regrettably, bioinformatics is not a magic wand. (Sadly, we are not bioinformagicians.) It cannot correct poor experimental design. It cannot overcome a lack of statistical power. Just like at the bench, if you design an experiment poorly, include confounding variables or overlook covariates, you might not be measuring what you think and/or you might not have any signal from your analysis. It is tempting to think that bioinformatics is more limited that bench science because we cannot collect our own data, but this would be wrong. Bench data is the raw material on which bioinformatics is performed. We can collect new data from other data - much of sequence analysis is doing just this. Of course, if our particular study focus of interest has no data, we need to generate it. But if you want to study the affect of a certain drug on a certain cell line and either/both do not exist, you have to generate that too.

So, what can we do about it? Bioinformaticians have to take some of the responsibility, largely because we are the ones that write software that perpetuates the myth that understanding parameters is not important. What do I mean by that? Well, often the documentation or “help” for bioinformatics tools is poorly written and poorly maintained - if it exists at all. When it does exist, it is usually written with expert users in mind. The novice is flooded with parameters and does not know which ones are important, or when. One solution is to write a series of protocols in the same vein as bench protocols, highlighting when one might want to change certain parameters - and which parameters are most important. (I am no saint in this department, sadly. If nothing else, this post has made me more determined to do better.)

The bottom line is quite simple, though:

Bioinformatics is science. Full stop. It is no better than other science. It is no worse than other science. People do it right. People do it wrong. However, if you are worried that it’s not real, the chances are that either you are doing it wrong, or you have deluded yourself about the “reality” of observations from bench science.

Saturday, 1 August 2015

The Day the Earth Smiled

This somehow passed me by when it happens (perhaps caught up with the upcoming move to Australia) but the latest episode of the Infinite Monkey Cage podcast (series 12, episode 4) featured a short segment on “The Day the Earth Smiled”.

From the Cassini Imaging website:

On July 19, 2013, in an event celebrated the world over, NASA’s Cassini spacecraft slipped into Saturn’s shadow and turned to image the planet, seven of its moons, its inner rings – and, in the background, our home planet, Earth.

The CICLOPS site has the full picture. This section is from the Wikipedia page and has Earth marked with an arrow.

Pretty humbling stuff.

There’s more at CICLOPS including a higher resolution image of the Earth and Moon:

Sunday, 12 July 2015

Developments in high throughput sequencing (June 2015 Edition)

This is nearly a month old now but Keith Bradnam’s ACGT blog a while back drew my attention to the June 2015 edition of Lex Nederbragt’s Developments in high throughput sequencing in which he plots Gigabases* per run against (log) read length (*the human genome is about 3Gb):

I’m particularly excited by the two technologies on the right of this graph, which represent the latest single molecule “long read” sequencing technologies, both of which we now have access to through the Ramaciotti Centre for Genomics. In fact, we got our first data from the PacBio RS II (right) and it’s looking good! (More on that later.)

Despite being a bioinformatician with a background in genetics, I have been keeping my distance a bit from “next generation sequencing” as the technical challenges of dealing with short read data far eclipse the scientific interest. (For me, that is - the kinds of things that I am most interested in do not suit short read data.) The new long read technologies are a real game changer, and I see a lot more genomics in my (and this blog’s) future.

Sunday, 21 June 2015

The importance of knowing how your data are scaled

A few weeks ago, there was a post on WEIT, The correlation between rejection of evolution and rejection of environmental regulation: what does it mean? It was triggered by a tweet about by the Washington post about a graph comparing attitudes to the environment and attitudes to evolution, broken down by religious affiliation:

We’ll get to the tweet later. First, the graph. It was from a US National Center for Science Education blog post based on 2007 data from the Pew Religious Landscape Study, examining two binary choice statements:

y-axis. Stricter environmental laws and regulations cost too many jobs and hurt the economy; or Stricter environmental laws and regulations are worth the cost.

x-axis. Evolution is the best explanation for the origins of human life on earth. (Agree/disagree)

Data was normalised onto a percentile scale with each circle representing (1) by position, the normalised percentile of that group’s response, (2) by area, the size of that group. (36,000 people were surveyed in total.)

The percentile normalisation method was based on a previous analysis of different Pew questions by Toby Grant, who explains it thus:

Geek note on measurement

The range of each dimension ranges from zero to 100. These scores were calculated by calculating the percentage of each religion giving each answer. The percentages were then subtracted (e.g., percent saying “smaller government” minus percent saying “bigger government”). The scores were then standardized using the mean and standard deviation for all of the scores. Finally, I converted the standardized scores into percentiles by mapping the standardized scores onto the standard Gaussian/normal distribution. The result is a score that represents the group’s average graded on the curve, literally.

A few things annoy me about this:

  1. This is not simply a “Geek note”. Knowing what was done to data is vital for understanding what a plot means. To be fair to Grant, he does mention that he is plotting percentiles in the graph legend. (As far as I can see, Robineau does not mention it anywhere!)
  2. By first normalising to the mean and then converting everything to percentiles, there is a double loss of quantitative information. Following the first normalisation, all you can do is compare groups - there is no absolute information about responses. Following the second, you cannot even compare the degree of difference. What this plot is basically doing is pulling in the outliers to make them look more similar to mean, and spreading out those similar to the mean to make them look more different.
  3. When converting to percentiles, the additional normalisations seem pointless. Unless I've misunderstood, if the data is truly normally distributed then the percentile of the fitted data should be the same as the percentile of the raw data. If not, you shouldn’t do the normalisation in the first place. Either way, I think you are just adding error and confusion. (There is no data presented to support the fact that these opinions are normally distributed.)

It is also worth noting that, to the unwary, the circle sizes could be misleading. The bigger the circle, the more data and the more accurate the estimation of the value. The small circles might have much more random sampling bias in their positions. (Under a null model where all groups are the same, you would expect the large circles to gravitate towards the mean, while the smaller circles should be the outliers.) Most importantly, circles that overlap are not more similar than circles that do not.

It would be more useful to have estimated standard errors plotted for each group. Again, because we have lost the quantitative information, we cannot tell whether a small difference in responses (possibly within measurement error) would have a big difference in percentiles. There are 36,000 people in total but some of the groups are less than 0.5% and therefore have fewer than 200 people.

Robineau’s plot uses the same method although he:

“didn’t rescale to the 0-100 scale, since I didn’t want this to seem like a percentage when it isn’t.”

It's not a percentage but it is a percentile, so 0-100 is entirely appropriate. Leaving it as -1.0 to +1.0 is in fact very misleading, as it implies that people are positive or negative with respect to the questions. In reality, positive just means “above average” and negative is “below average”. I have an above average number of arms: two. This does not mean that I have lots of arms, it just means that some people have fewer arms than me.

These things aside, Robineau asks:

“So what does this tell us?

Thanks to the scaling, the only thing this graph tells us is that (a) there is a rank correlation between the answers to the two questions, and (b) some religious groups (particularly evangelical Christians) appear to agree with these statements less than average, while other groups (notably non-Christians) tend to agree with these statements more than average.

These observations could still be of interest. The real problem comes when people start interpreting this graph as if the normalisations and rescaling have not been done to it. Robineau first:

First, look at all those groups whose members support evolution. There are way more of them than there are of the creationist groups, and those circles are bigger. We need to get more of the pro-evolution religious out of the closet.

Second, look at all those religious groups whose members support climate change action. Catholics fall a bit below the zero line on average, but I have to suspect that the forthcoming papal encyclical on the environment will shake that up.”

This in turn was apparently interpreted by the Washington post to mean this:

The fact is, the normalisation has removed all hope of actually knowing whether there is conflict or not. The percentile scaling removes almost all of the quantitative info on the axes, so proximity on the scale means nothing with respect to proximity of answer. All the groups inside the small top right cluster could have >90% support for the scientific evidence and all of the groups outside <10% support, and you could still get that plot. (It’s hard to tell but the top-right cluster look closer to 1.0 than the bottom-left groups are to -1.0, indicating that they might deviate much more from the mean thanks to the mapping onto a normal distribution. This implies that the data was not normally distributed in the first place and is probably a heavy-tailed or bimodal distribution instead.)

Critically, it is impossible to conclude that any groups “support evolution” or “support climate change action”. As the graph is scaled by percentiles, 0.0 is essentially the point where 50% are above and 50% below. Because the vast majority of groups are religious, of course there are many religious groups above the line. There essentially have to be, unless all religious groups were identical (in which case they would group very slightly below 0.0).

To many, stand-out thing is that atheists and agnostics are all in the top-top right. This graph could easily have been branded “the conflict between science and religion in one chart”! But it cannot even really say that: every group could disagree with the two statements and thus be in conflict with the scientific evidence. You would still get the same plot after the rescaling.

My big question from all of this is: why not make the plot using the raw percentage responses? What do the normalisations actually achieve?

And my big take home message: if you are going to infer things from plots, make sure that you understand how the data were scaled.

Tuesday, 28 April 2015

May's SCB Conservation Cafe is all about Herpetofauna

The Sydney Society for Conservation Biology (SCB) have Michael McFadden, the Unit Supervisor of the Herpetofauna division at Taronga Zoo, for 2nd May’s Conservation Cafe. That’s reptiles and amphibians to the rest of us:

This May, Sydney-SCB welcomes Michael McFadden, the Unit Supervisor of the Herpetofauna division at Taronga Zoo. Michael began working at Taronga Zoo in January 2003 and now oversees the maintenance and husbandry of the Zoo’s collection of reptiles and amphibians. He works closely with the Zoo’s conservation projects which include captive breeding and release programs for the highly endangered Southern and Northern Corroboree Frogs. The current focus of Michael’s work is developing techniques to improve captive breeding and rearing success in threatened Australian frogs and reintroduction biology.

As before, it’s free: RSVP on Eventbrite.

Monday, 27 April 2015

Yet another study debunks yet another vaccine/autism myth

As reported by ABC Science last week, a large study (roughly 95,000 people) has hammered another nail into the well-and-truly debunked vaccine-autism link.

In an accompanying editorial in [the Journal of the American Medical Association (JAMA)], Dr Bryan King, a doctor at the University of Washington and Seattle Children’s Hospital, says the data is clear.

“The only conclusion that can be drawn from the study is that there is no signal to suggest a relationship between MMR and the development of autism in children with or without a sibling who has autism,” writes King.

“Taken together, some dozen studies have now shown that the age of onset of ASD does not differ between vaccinated and unvaccinated children, the severity or course of ASD does not differ between vaccinated and unvaccinated children, and now the risk of ASD recurrence in families does not differ between vaccinated and unvaccinated children.”

Make no mistake: if you avoid vaccines in fear of autism, you are a misguided fool. If you spread this myth, you are an ignorant menace to society. This may sound harsh but people can really die if anti-vaxxers get their way.

Reference:

Jain A, Marshall J, Buikem A, Bancroft T, Kelly JP & Newschaffer CJ (2015): Autism Occurrence by MMR Vaccine Status Among US Children With Older Siblings With and Without Autism. JAMA 313(15): 1534-1540.

Sunday, 26 April 2015

Best journal cover ever?

Courtesy of the Molecular Biology and Evolution Facebook page comes this awesome cover art:

According to the MBE Editor:

The author and artist info: The cover image depicts representative squamate species (lizards and snakes) playing poker, with the card and chip colors representing the sex-determining system most prevalent in each clade. The tabletop shows results from a comparative genomic analysis of squamate sex-determining mechanisms by Gamble et al in this issue. This study discovered that changes between sex-determining mechanisms in one clade, geckos, account for a half to two-thirds of the total transitions known in lizards and snakes. This remarkable frequency of transition is reflected in the illustration by the heightened activity at the gecko side of the table: the three gecko species in the foreground are cheating, implying that when it comes to sex determination, geckos do not play by the rules. The image was created by University of Minnesota biologist and artist Anna Minkina and pays homage to the Cassius M. Coolidge painting, “A Friend in Need”, part of the artist’s “Dogs Playing Poker” series.

h/t: James McInereny

Sunday, 19 April 2015

Good news! WHO calls for results from all trials to be reported

A bit belated but good news worth sharing nonetheless. From Ian Bushfield at Sense About Science, and as reported in Science and other media sites:

For the first time ever, the World Health Organisation (WHO) has taken a position on clinical trial results reporting, and it’s a very strong position! The WHO now says that researchers have a clear ethical duty to publicly report the results of all clinical trials. Significantly, the WHO has stressed the need to make results from previously hidden trials available. Ben Goldacre said, “This is a very positive, clear statement from WHO, and it is very welcome.” Ilaria Passarani from the European Consumer Organisation BEUC called it “a landmark move for consumers.” It is the position we and hundreds of you wrote to the WHO last autumn urging them to adopt. Well done everyone!

You can read more about the WHO’s statement and responses to it on the AllTrials website.

Further reading: Goldacre B (2005): How to Get All Trials Reported: Audit, Better Data, and Individual Accountability. PLoS Medicine 12(4): e1001821.

Tuesday, 24 March 2015

Yet more OMICS Group spam full of vacuous rubbish

Despite unsubscribing from a mailing list that I never subscribed to, OMICS Group keep sending me pointless emails to crappy conferences. Today it was “Biodiversity-2015”:

Biodiversity-2015 is specifically premeditated with a unifying axiom providing pulpit to widen the imminent scientific creations. The main theme of the conference is “Share and Enhance Ecological & Geological Conservation research” which covers a broad array of vitally key sessions.

“Biodiversity-2015 is specifically premeditated with a unifying axiom providing pulpit to widen the imminent scientific creations.” Wow! Someone had been over-using their random vacuous crap generator.

Again, no explicit mention of OMICS Group as the organiser was made, although this one did mention “accepted abstracts will be published in the respective OMICS Group Journals”. (For free - they’re so generous!)

At the end of the email, they tell me to:

Have a Great Day Doctor!!

Well, with two exclamation marks, do I have any choice?! What would have made a greater day would have been (a) not receiving the email in the first place*, and (b) having the “To unsubscribe click here” line at the end of the email actually contain a hyperlink. Instead, the “click here” was just text in a different colour!

*I’m not being entirely honest with (a) - I think my day was brightened a little by “specifically premeditated with a unifying axiom providing pulpit to widen the imminent scientific creations”!

Sunday, 22 March 2015

OMICS Group strike again with more scam conference spam

I am always wary when I receive an email that begins something to the effect of:

Dear Colleague,

The purpose of this letter is to solicit your gracious presence as a Speaker at the upcoming 5th World Congress on Cancer Therapy on September 28-30, 2015 which is going to be held in Atlanta, USA.

Soliciting my gracious presence smacked of an OMICS Group “predatory” conference invitation. The rest of the email went on:

The aim of this conference is to learn and share knowledge in cancer research. Leading World cancer researchers, Public health professionals, scientists, academic scientists, World Breast Surgeons, Medical and Surgical Oncologists, Radiologists, Researchers, Healthcare professionals, Industry researchers, Nurses, Scholars, Decision makers, Students and other professionals gather in Atlanta to speak at our conference.

Exceptional Benefits

All accepted abstracts will be published in the respective Journals
Each abstract will receive a DOI provided by Cross Ref
Certification by the organizing committee
Global Exposure to your Research
Best Poster Competitions and Young Researcher Competitions
The Career Guidance Workshops to the Graduates Doctorates and Post-Doctoral Fellows
Networking with Experts across the globe

For more details on scientific sessions and abstract submission, please Click Here

In closing, we would be pleased and honored if you would consent to be our speaker at our Conference.

I will call you in a week or so to follow up on this.

Regards,
Isaac Bruce
Cancer Therapy 2015
2360 Corporate Circle
Suite 400 Henderson
NV 89074-7722, USA
cancertherapy@conferenceseries.com

The thing I find most curious is that there is not direct reference to OMICS Group anywhere in this information: they don’t even name the “respective Journals”, which are presumably OMICS journals as with their other conferences. The address and even the email address are equally opaque.

If you do “Click Here” then you go to an abstract submission page with the “OMICS International” logo and some OMICS Group references/email addresses but even here they opt not to use an omicsgroup.com URL.

This does not strike me as the behaviour of an organisation that is proud of their brand. Indeed, I suspect that they know that their brand is toxic thanks to their reputation for predatory journals and conferences and thus try to lure people to submit an abstract (with a hefty $899 registration fee) before they realise their mistake.

The only other clue was the line hidden at the bottom in small font:

You are subscribed to OMICS Group as XXX. If you do not wish to receive any further communications, please click here.

Suffice it to say that I clicked there.

Thursday, 22 January 2015

Antibiotics really matter and need more research funding

Last year, I went down with tonsillitis on New Year’s Eve. The timing sucked a bit but I guess there are no good times for such things to happen. However, I was/am lucky: lucky because I live in an age where antibiotics are available and still work.

As a scientist, I often get frustrated about science funding. Firstly, there’s not enough of it (compared to other endeavours of less benefit to both society and the economy) but secondly, a lot of it goes to the wrong places, namely human diseases such as cancer. Don’t get me wrong: I’d love to see cancer cured. It’s just that there are bigger fish to fry, and global crises looming that would make diseases of old age such as cancer and Alzheimer’s a bit of a moot point.

An obvious need of greater funding is climate change, thrown into the spotlight again (as if it were needed) by the confirmation that 2014 was the the warmest year on record. Another is the development of new antibiotics.

Antibiotic resistance is a big problem and one that is currently only getting worse. We are heading for a “post antibiotic world”, which is a really scary thought. Indeed, some scientists have argued that antibiotic resistance is a bigger problem than climate change because we have the technology to combat climate change, we just lack the political will. We do not yet have the technical solution to the impending “antibiotic apocalypse”, hence the real need to throw money at research into solutions.

The annoying thing is that this is not a problem that has snuck up on us. In an editorial from 1997 entitled “Antibiotic Armageddon”, Calvin Kunin from The Ohio State University wrote:

“The advances of the antimicrobial era are being dissipated by the emergence and spread of resistant microorganisms, the inevitable consequence of intense use of antibiotics in humans over the past 50 years. The process is accelerating in the community as well as in hospitals and is a problem worldwide. The attrition of older drugs is sustained by the selective effects of new and more expensive drugs developed to overcome resistance. Novel compounds will no doubt be discovered, but their demise is inevitable. It is just a matter of time until resistant pyogenic organisms join the opportunistic microbes as major threats to humans.”

…

“[The] long-term outlook for control of antibiotic resistance is bleak. There are simply too many physicians prescribing antibiotics casually and too many people buying antibiotics without a prescription in developing countries. There is only a thin red line of infectious diseases practitioners who have dedicated themselves to rational therapy and control of hospital infections. The issues need to be presented forcefully to the medical community and the public. Third-party payers must get the message that these programs can save lives as well as money.”

Sadly, if these words were written today, I don’t think anyone would argue the point; not much has changed. And that’s without even mentioning the big problems caused by the long-running over-use of antibiotics is agriculture.

Things are not without hope. Earlier this month, a Nature paper by Lin et al. reported the discovery of a novel class of antibiotic from a screen of 10,000 bacterial strains, following the development of novel method to grow hitherto uncultured bacteria:

“Antibiotic resistance is spreading faster than the introduction of new compounds into clinical practice, causing a public health crisis. Most antibiotics were produced by screening soil microorganisms, but this limited resource of cultivable bacteria was overmined by the 1960s. Synthetic approaches to produce antibiotics have been unable to replace this platform. Uncultured bacteria make up approximately 99% of all species in external environments, and are an untapped source of new antibiotics. We developed several methods to grow uncultured organisms by cultivation in situ or by using specific growth factors. Here we report a new antibiotic that we term teixobactin, discovered in a screen of uncultured bacteria. Teixobactin inhibits cell wall synthesis by binding to a highly conserved motif of lipid II (precursor of peptidoglycan) and lipid III (precursor of cell wall teichoic acid). We did not obtain any mutants of Staphylococcus aureus or Mycobacterium tuberculosis resistant to teixobactin. The properties of this compound suggest a path towards developing antibiotics that are likely to avoid development of resistance.”

Personally, I remain skeptical about claims that teixobactin is resistance-proof. It may not be easy but I am sure the bugs will stumble across a way to evade or destroy the toxin and evolve resistance. Nonetheless, the message is clear: there are new antibiotics out there to be found. This one was found in the backyard of one of the researchers! Bacteria have been killing each other for millions, maybe billions, of years and so the global diversity in nature is likely to be massive. We just need the ingenuity and funding to find them.

References

Farrar J & Woolhouse M (2014). Policy: An intergovernmental panel on antimicrobial resistance. Nature 509:555–557

Kunin CM (1997). Antibiotic Armageddon. Clinical Infectious Diseases 25:240–1

Lin LL et al. (2015). A new antibiotic kills pathogens without detectable resistance. Nature doi:10.1038/nature14098