On 2014 Jul 16, Claudiu Bandea commented:
Closing the gap between ‘words’ and ‘facts’ in evaluating genome biology and the ENCODE project
When I referred to Ford Doolittle’s article (1) as “7 pages of small print text” (2), I meant it both, literary and figuratively. Indeed, Doolittle’s paper is a remarkable example of fine print, which is likely to induce heated arguments for a long time, just as the author concluded: “many of the most heated arguments in biology are not about facts at all but rather about the words that we use to describe what we think the facts might be.”
Remarkably, more than 40 years ago, Susumu Ohno started his famous paper on junk DNA (3) with a related statement: “Over the years, I have learned that there is no such thing as a fact. What passes for a fact is in truth a set of observations and its interpretation. Therefore, the interpretation is just as important to a fact as the observation itself.”
In my previous mini-essay (2) on Doolittle’s article (1), I made 3 main points:
(i) Ever since the notion of ‘junk DNA’ (jDNA) was introduced as a metaphor for presumably non-functional genomic DNA in species with relatively high C-value, it has been clearly used by the scientists in the field of genome biology and evolution to denote genomic DNA that has no biological function at all, whether informational (iDNA) or non-informational (niDNA), and implying otherwise is nonsensical.
(ii) The two main theories on putative non-informational functions for the so called jDNA, the nucleo-skeletal and the nucleotypic functions, which “describe genome size variation as the outcome of selection via the intermediate of cell size” (4), and which Doolittle uses as pillars for his theoretical framework on genome evolution and biology, do not explain the C-value paradox.
(iii) Apparently, Doolittle is not aware of the theory that most genomic niDNA, redefined as symbiotic DNA (sDNA), functions as a protective mechanism (adaptive genomic immunity) against deleterious insertional mutagenesis by endogenous and exogenous inserting elements, such as retroviruses, and that this theory is fully supported by the current data and observations and it explains the C-value paradox (5, 6).
Here, I attempt to further close the gap between 'words' and 'facts' in addressing the genome biology and in evaluating the ENCODE project.
In the introductory section, Doolittle outlines the premise for his paper: “a flurry of articles and letters”, published in Nature in other journals under the umbrella of the ENCODE project, “collectively claim function for the majority of the 3.2 Gb human genome”, which, if true, would debunk the notion of jDNA (1). The problem with Doolittle’s premise, however, is that it is not based on facts; indeed, the “flurry of ENCODE publications” did not claim that the majority of the human genome is functional (7). On the contrary, in what seems to have been a concerted, but tacit ‘silence policy,’ the ENCODE authors went out of their way not to address the ‘functionally’ of the human genome in their publications. In light of this fact, Doolittle had no choice but to build the premise for his paper on secondary sources offered by various science writers (8-10), who were apparently caught into a publicity stunt orchestrated on the side by a few of ENCODE scientists. Whether Doolittle’s approach of using secondary sources, which is a strong departure from conventional academic standards, sets up a hasty precedent for the scientific literature remains to be seen.
So, why wasn’t the ENCODE project designed in context of the fundamental issues and knowledge about genome biology and evolution, such as the C-value paradox, limited sequence conservation among closely related species, mutational load, and the evolutionary origin of most genomic sequences from transposable elements? And, why did the ENCODE researchers choose not to address these fundamental issues in their official publications? Obviously, this makes no sense considering that their massive and expensive project was funded specifically to annotate the ‘functional sequences’ of the human genome.
Fully addressing these questions might take us deep into the science of human behavior, and might highlight deep deficiencies in our current system of funding science, which relies on a weak and closed peer review system (parenthetically, a sensible solution would be to replace this system, which is vulnerable to abuse, with a stronger and true peer review system that is open to all peers).
Nevertheless, it is inconceivable that the ENCODE leaders, who represent some of the finest science institutions in the world, were scientifically incompetent, as suggested by some critics of ENCODE (10), and were not aware of these fundamental issues on genome biology and evolution. On the contrary, it was the appreciation for this fundamental knowledge that prompted them to be silent, as this knowledge is in conflict with some of their study objectives and raises inconvenient questions about the relevance of their study and results.
However, as illustrated by Doolittle’s article, the full significance of this fundamental knowledge on genome biology and evolution is not clearly recognized in the field, which has led to tremendous confusion and has allowed projects such as ENCODE to flourish. Indeed, unfortunately, the knowledge on genome biology and evolution has yet to crystalize in clear facts. However, here is one (in large print): based on the C-value paradox, limited sequence conservation, mutational load, and the evolutionary origin of most genomic sequences from transposable elements, it is clear that MOST OF THE HUMAN GENOME CANNOT HAVE INFORMATIONAL FUNCTIONS, period.
Now that we have cracked ENCODE’s ‘code of silence’, reset some of Doolittle’s small print, and crystalized the fact that only a small fraction of the human genome has informational functions, it is time to focus on the major question remaining in the field of genome evolution and biology:
Does most of the genome in organisms with relatively high C-value have non-informational functions, or most of it is non-functional, metaphorically speaking junk?
References
(1) Doolittle WF. 2013. Is junk DNA bunk? A critique of ENCODE. Proc Natl Acad Sci USA., 110:5294-300. Doolittle WF, 2013
(2) Bandea CI. 2013. Junk DNA is bunk, but not as suggested by ENCODE or Doolittle. PubMed Commons (National Library of Medicine; Bethesda, MD). Comment on: Doolittle WF, 2013
(3) Ohno S. 1973. Evolutional reason for having so much junk DNA. In Modern Aspects of Cytogenetics: Constitutive Heterochromatin in Man (ed. R.A. Pfeiffer), pp. 169-173. F.K. Schattauer Verlag, Stuttgart, Germany.
(4) Gregory TR. 2004. Insertion-deletion biases and the evolution of genome size. Gene, 324:15-34. Gregory TR, 2004
(5) Bandea CI. 1990. A protective function for noncoding, or secondary DNA. Med. Hypoth., 31:33-4. Bandea CI, 1990
(6) Bandea CI. 2013. On the concept of biological function, junk DNA and the gospels of ENCODE and Graur et al. bioRxiv doi: 10.1101/000588; http://biorxiv.org/content/early/2013/11/18/000588
(7) ENCODE Project Consortium. 2012. An integrated encyclopedia of DNA elements in the human genome. Nature, 489:57–74. ENCODE Project Consortium., 2012
(8) Kolata G. 2012 (September 5). Bits of mystery DNA, far from ‘junk’, play crucial role. The New York Times, Section A, p. 1.
(9) Anonymous, 2012. Cracking ENCODE. Lancet, 380:950. Anonymous, 2012
(10) Pennisi E. 2012. Genomics. ENCODE project writes eulogy for junk DNA. Science, 337:1159–1161. Pennisi E, 2012
(11) Graur D et al., 2013. On the immortality of television sets: "function" in the human genome according to the evolution-free gospel of ENCODE. Genome Biol Evol., 5(3):578-90.Graur D, 2013
This comment, imported by Hypothesis from PubMed Commons, is licensed under CC BY.