Molecular Biology and Evolution

Discovery of Novel Genes Derived from Transposable Elements Using Integrative Genomic Analysis

Hoen, D. R., Bureau, T. E..

Complex eukaryotes contain millions of transposable elements (TEs), comprising large fractions of their nuclear genomes. TEs consist of structural, regulatory, and coding sequences that are ordinarily associated with transposition, but that occasionally confer on the organism a selective advantage and may thereby become exapted. Exapted transposable element genes (ETEs) are known to play critical roles in diverse systems, from vertebrate adaptive immunity to plant development. Yet despite their evident importance, most ETEs have been identified fortuitously and few systematic searches have been conducted, suggesting that additional ETEs may await discovery. To explore this possibility, we develop a comprehensive systematic approach to searching for ETEs. We use TE-specific conserved domains to identify with high precision genes derived from TEs and screen them for signatures of exaptation based on their similarities to reference sets of known ETEs, conventional (non-TE) genes, and TE genes across diverse genetic attributes including repetitiveness, conservation of genomic location and sequence, and levels of expression and repressive small RNAs. Applying this approach in the model plant Arabidopsis thaliana, we discover a surprisingly large number of novel high confidence ETEs. Intriguingly, unlike known plant ETEs, several of the novel ETE families form tandemly arrayed gene clusters, whereas others are relatively young. Our results not only identify novel TE-derived genes that may have practical applications but also challenge the notion that TE exaptation is merely a relic of ancient life, instead suggesting that it may continue to fundamentally drive evolution.