Butterfly genome reveals promiscuous exchange of mimicry adaptations among species.
The evolutionary importance of hybridization and introgression has long been debated1. We used genomic tools to investigate introgression in Heliconius, a rapidly radiating genus of neotropical butterflies widely used in studies of ecology, behaviour, mimicry and speciation2-5 . We sequenced the genome of Heliconius melpomene and compared it with other taxa to investigate chromosomal evolution in Lepidoptera and gene flow among multiple Heliconius species and races. Among 12,657 predicted genes for Heliconius, biologically important expansions of families of chemosensory and Hox genes are particularly noteworthy. Chromosomal organisation has remained broadly conserved since the Cretaceous, when butterflies split from the silkmoth lineage. Using genomic resequencing, we show hybrid exchange of genes between three co-mimics, H. melpomene, H. timareta, and H. elevatus, especially at two genomic regions that control mimicry pattern. Closely related Heliconius species clearly exchange protective colour pattern genes promiscuously, implying a major role for hybridization in adaptive radiation.
- Research Article
148
- 10.1186/s13059-016-0889-0
- Feb 27, 2016
- Genome Biology
BackgroundAlthough hybridization is thought to be relatively rare in animals, the raw genetic material introduced via introgression may play an important role in fueling adaptation and adaptive radiation. The butterfly genus Heliconius is an excellent system to study hybridization and introgression but most studies have focused on closely related species such as H. cydno and H. melpomene. Here we characterize genome-wide patterns of introgression between H. besckei, the only species with a red and yellow banded ‘postman’ wing pattern in the tiger-striped silvaniform clade, and co-mimetic H. melpomene nanna.ResultsWe find a pronounced signature of putative introgression from H. melpomene into H. besckei in the genomic region upstream of the gene optix, known to control red wing patterning, suggesting adaptive introgression of wing pattern mimicry between these two distantly related species. At least 39 additional genomic regions show signals of introgression as strong or stronger than this mimicry locus. Gene flow has been on-going, with evidence of gene exchange at multiple time points, and bidirectional, moving from the melpomene to the silvaniform clade and vice versa. The history of gene exchange has also been complex, with contributions from multiple silvaniform species in addition to H. besckei. We also detect a signature of ancient introgression of the entire Z chromosome between the silvaniform and melpomene/cydno clades.ConclusionsOur study provides a genome-wide portrait of introgression between distantly related butterfly species. We further propose a comprehensive and efficient workflow for gene flow identification in genomic data sets.Electronic supplementary materialThe online version of this article (doi:10.1186/s13059-016-0889-0) contains supplementary material, which is available to authorized users.
- Research Article
3
- 10.1111/j.1365-294x.2007.03660.x
- Jan 1, 2008
- Molecular Ecology
Editorial and Retrospective 2008
- Research Article
38
- 10.2307/2446495
- Nov 1, 1998
- American Journal of Botany
Reappraising adaptive radiation
- Research Article
8
- 10.1186/s12864-023-09826-z
- Nov 24, 2023
- BMC Genomics
BackgroundBryozoans are mostly sessile aquatic colonial invertebrates belonging to the clade Lophotrochozoa, which unites many protostome bilaterian phyla such as molluscs, annelids and brachiopods. While Hox and ParaHox genes have been extensively studied in various lophotrochozoan lineages, investigations on Hox and ParaHox gene complements in bryozoans are scarce.ResultsHerein, we present the most comprehensive survey of Hox and ParaHox gene complements in bryozoans using four genomes and 35 transcriptomes representing all bryozoan clades: Cheilostomata, Ctenostomata, Cyclostomata and Phylactolaemata. Using similarity searches, phylogenetic analyses and detailed manual curation, we have identified five Hox genes in bryozoans (pb, Dfd, Lox5, Lox4 and Post2) and one ParaHox gene (Cdx). Interestingly, we observed lineage-specific duplication of certain Hox and ParaHox genes (Dfd, Lox5 and Cdx) in some bryozoan lineages.ConclusionsThe bryozoan Hox cluster does not retain the ancestral lophotrochozoan condition but appears relatively simple (includes only five genes) and broken into two genomic regions, characterized by the loss and duplication of serval genes. Importantly, bryozoans share the lack of two Hox genes (Post1 and Scr) with their proposed sister-taxon, Phoronida, which suggests that those genes were missing in the most common ancestor of bryozoans and phoronids.
- Research Article
12
- 10.1186/s12862-016-0692-2
- Jun 16, 2016
- BMC Evolutionary Biology
BackgroundUnderstanding the mechanisms and selective forces leading to adaptive radiations and origin of biodiversity is a major goal of evolutionary biology. Acrocephalus warblers are small passerines that underwent an adaptive radiation in the last approximately 10 million years that gave rise to 37 extant species, many of which still hybridize in nature. Acrocephalus warblers have served as model organisms for a wide variety of ecological and behavioral studies, yet our knowledge of mechanisms and selective forces driving their radiation is limited. Here we studied patterns of interspecific gene flow and selection across three European Acrocephalus warblers to get a first insight into mechanisms of radiation of this avian group.ResultsWe analyzed nucleotide variation at eight nuclear loci in three hybridizing Acrocephalus species with overlapping breeding ranges in Europe. Using an isolation-with-migration model for multiple populations, we found evidence for unidirectional gene flow from A. scirpaceus to A. palustris and from A. palustris to A. dumetorum. Gene flow was higher between genetically more closely related A. scirpaceus and A. palustris than between ecologically more similar A. palustris and A. dumetorum, suggesting that gradual accumulation of intrinsic barriers rather than divergent ecological selection are more efficient in restricting interspecific gene flow in Acrocephalus warblers. Although levels of genetic differentiation between different species pairs were in general not correlated, we found signatures of apparently independent instances of positive selection at the same two Z-linked loci in multiple species.ConclusionsOur study brings the first evidence that gene flow occurred during Acrocephalus radiation and not only between sister species. Interspecific gene flow could thus be an important source of genetic variation in individual Acrocephalus species and could have accelerated adaptive evolution and speciation rate in this avian group by creating novel genetic combinations and new phenotypes. Independent instances of positive selection at the same loci in multiple species indicate an interesting possibility that the same loci might have contributed to reproductive isolation in several speciation events.Electronic supplementary materialThe online version of this article (doi:10.1186/s12862-016-0692-2) contains supplementary material, which is available to authorized users.
- Research Article
12
- 10.2307/3558369
- Oct 1, 2001
- American Journal of Botany
Given these remarkable developments, and the three decades that have elapsed since the last comprehensive-and magisterial-treatments of plant speciation and evolution by Levin was a hero of my youth, who had made his name in plant population biology, but also added substantially to our understanding of speciation through his research on the genus Phlox (Polemoniaceae). His contributions included early studies of plant demography, genetic differentiation among populations, divergence in pollination biology among closely related species, selection for reproductive isolation, and constraints on species ranges.
- Book Chapter
1
- 10.1007/978-1-4615-1929-4_1
- Jan 1, 1995
HOM/Hox genes are a family of genes (the homeogene family) that show similarities in structure, organisation into complexes and expression patterns. They code for transcriptional regulators which are thought to function at the top of a genetic hierarchy that controls the relative position of the cells along the embryonic axes. First discovered in Drosophila and then found in vertebrates, they were recently identified in the genome of animals as diverse as nematodes, leeches, amphioxus and hydra, suggesting that a HOM/Hox gene cluster already existed in the common ancestor of all animals. The presence of HOM/Hox cluster(s) has been proposed as one of the characters defining the Kingdom Animalia (i.e. the Zootype: reviewed in Slack, 1993; Holland, 1992; Thorogood, 1993; Kappen et al., 1993). Interphyletic comparisons between insect HOM genes and vertebrate Hox genes are well documented (reviewed in McGinnis and Krumlauf, 1992; Botas, 1993). In vertebrates there are four paralogous clusters of Hox genes, referred to as Hox A, B, C, and D, each showing clear structural homology to the prototypic homeotic complex (HOM-C) of Drosophila. During ontogenesis, Hox genes are expressed in the neurectoderm and paraxial mesoderm in specific but overlapping domains that extend from the caudal end of the embryo to a sharp rostral limit. There is a correlation (termed spatial colinearity) between the position of a Hox gene within its cluster and the location of its rostral limit of expression. Significantly, this spatial colinearity also exists in Drosophila. Both vertebrate Hox and insect HOM gene expression domains respect metameric boundaries. Loss-of-function mutations (i.e. generated by gene disruption) and gain-of-function mutations (i.e. generated by ectopic gene expression) of both mouse Hox genes and Drosophila HOM genes often cause a transformation of specific metameres [serially homologous anatomical units, repeated along the rostrocaudal axis such as somites (mouse) or parasegments (Drosophila)] into the likeness of their neighbours, indicating that these genes are involved in the specification of the phenotype of a given segment according to its position along the rostrocaudal axis. Since the segmented body plans of mouse and Drosophila have arisen totally independently during evolution, it appears that at least part of HOM/Hox genes network has been co-opted to impart morphogenetic segment identity in animal groups employing different segmentation strategies (Holland, 1990).
- Research Article
3
- 10.1038/sj.embor.7400804
- Aug 1, 2006
- EMBO reports
ThisThis workshop brought together developmental biologists and oncologists for three intense days of discussions about the recent progress that has been made on the roles of homeodomain proteins in development, haematopoiesis and leukaemogenesis. The homeo-domain is a 61-amino-acid DNA-binding motif with specific sequence and structural characteristics (Gehring & Hiromi, 1986). Homeodomain proteins were initially identified in Drosophila as being responsible, when mutated, for homeotic transformations—that is, the conversion of a part of the body into the likeness of another. Since their initial discovery, the number of homeodomain proteins has increased enormously, and they comprise several different families of transcription factors. The main stars of the meeting were members of the Hox family and of the three amino-acid loop extension (TALE) families, with some other homeodomain proteins making guest appearances. TALE proteins include the PBC family (Exd in Drosophila and Pbx proteins in vertebrates) and the Hth/Meis superfamily (Hth in Drosophila, and Meis and Prep proteins in vertebrates). The link between Hox and TALE proteins is their functional interactions, which were initially revealed in flies by genetic analyses and later supported by biochemical data (Mann & Affolter, 1998; Moens & Selleri, 2006). PBC proteins have been shown to act as Hox cofactors that mainly determine Hox DNA-binding specificity and selectivity. Hth/Meis proteins interact with PBC proteins, participating in ternary complexes with Hox proteins, or modulating PBC protein stability and subcellular localization. This European Molecular Biology Organization/European School of Molecular Medicine workshop on Homeodomain Proteins, Hematopoietic Development and Leukemias took place in Riva del Garda, Italy, between 23 and 25 March 2006, and was organized by F. Blasi, ... Functional specificities and redundancies A large part of the meeting was devoted to the presentation and discussion of the consequences of the complete or partial inactivation of different TALE products in several vertebrates. These studies arrived at two main conclusions: first, TALE proteins are essential for the development of most areas of the organism; and second, they have both individual functions and a large degree of redundancy. This was clearly shown in the Pbx gene family, for which an advanced phenotypic analysis of single and compound mutants was presented. In mice, the individual contribution of each member of the family seemed to be different because, whereas the Pbx1 single mutant had a wide range of malformations, only minor or no phenotypes were obvious in Pbx2 and Pbx3 mutants. However, when a reduced Pbx2 and/or Pbx3 dosage was added to the Pbx1−/− background, strong exacerbations of the malformations were detected, as reported by L. Selleri (New York, NY, USA) and M. Cleary (Stanford, CA, USA). In particular, the cardiovascular system, the axial skeleton, the limbs and the craniofacial area were shown to be strongly affected. A.J. Waskiewicz (Edmonton, Canada) also described redundancy among Pbx genes in zebrafish, in which the functions of the Pbx2 and Pbx4 gene products seem to overlap in a range of tissues, including the hindbrain and tectum, and during early haematopoiesis. Hox gene products are known to be largely redundant, especially within paralogous groups. In the case of the Meis and Prep gene families, pleiotropic functions in patterning and differentiation processes were also described (see below); however, genetic analyses designed to evaluate the possible extent of redundancy among the different members of these gene families are still at an early stage. At the meeting, several tissues were reported to require the activity of homeodomain proteins for proper development; we highlight some of these data below.
- Research Article
20
- 10.1111/j.1600-0706.2009.17835.x
- Aug 27, 2009
- Oikos
Which ecologically important traits are most likely to evolve rapidly?
- Research Article
15
- 10.1111/nph.13826
- Dec 22, 2015
- New Phytologist
Opportunities for unlocking the potential of genomics for African trees.
- Research Article
4
- 10.1095/biolreprod.112.101402
- Jun 1, 2012
- Biology of Reproduction
Homeobox (Hox) genes encode several subfamilies of nuclear proteins with structurally conserved helix-turn-helix domains that recognize specific DNA sequences [1]. As transcription factors, HOX family members control key events during early development, influencing the organization of body patterning and tissue structures in all species of the animal kingdom from invertebrates to humans [1, 2]. In fully developed organisms, Hox genes contribute to tissue damage and repair mechanisms by controlling apoptosis, receptor signaling, differentiation, motility, and angiogenesis; when expressed abnormally, these activities can have oncogenic consequences [3]. Rhox5 was first identified by Wilkinson et al. [4] as an oncofetal gene (originally called Pem because it was expressed in placenta and embryos) that influenced the progression of Tcell lymphomas to their oncogenic phenotype. During the decade following its discovery, Pem was classified as an orphan member of the Hox gene family that is selectively expressed in reproductive tissues. It therefore followed that Pem was renamed Rhox5 when it was located within a syntenically conserved cluster of other, related reproductive Hox genes on the mouse X chromosome, which was dubbed the Rhox gene family locus in a groundbreaking report in 2005 [5]. In the intervening years, MacLean, Wilkinson, and their colleagues defined the roles of Rhox genes in reproductive tissues [6]. In an elegant report published in the current issue of Biology of Reproduction, MacLean et al. [7] have integrated the power of mouse genetics with the emerging discipline of comparative genomics to gain insight into how Rhox gene expression is regulated within specific regions of the epididymis. The rodent epididymis is ideal for these studies because of its well-established segmentation into three anatomically and functionally distinct regions; the head, or caput; the main body, or corpus; and the tail, or cauda. Over the ;16M years since the Mus/Rattus evolutionary divergence, the segmented anatomical arrangement of the rodent epididymis has remained intact, but each of the segments has acquired distinct roles in terms of the maturation and storage of spermatozoa in the different species. In a clear example of this, the caput epididymidis of mice plays a key role in sperm maturation and the acquisition of forward motility, whereas this occurs in the rat cauda epididymidis, which is normally considered to be a site of sperm storage. Building on a previous finding that the sperm of Rhox5 null mice lack forward motility [5], MacLean et al. [7] embarked on a series of experiments. Their results now show that Rhox5 is a master regulator of many other members of the X-linked Rhox gene cluster in the mouse caput epididymidis and that it acts in almost the same way in the rat cauda epididymidis. This remarkable evolutionary shift in the Rhox5-dominated regionspecific expression of Rhox genes in the rodent epididymis sets the stage for additional experiments that promise to provide insight into mechanisms responsible for the functional maturation of sperm. Comparative studies of gene expression profiles of different segments of the rat and mouse epididymis have already been performed [8], and similar studies will undoubtedly follow in mice with either a functional or a disrupted Rhox5 gene. This should help define whether Rhox5 or other members of the Rhox family act directly or indirectly by controlling spatially defined networks of genes to explain why sperm from the mouse caput and corpus epididymidis are functionally much more mature than rat sperm taken from the same anatomical locations. These types of studies integrate the power of comparative genomics with a wealth of information about the functional diversity of reproductive tissues in two closely related species, and they offer huge potential for the discovery of key mechanisms that control male fertility. In this context, the report by MacLean et al. [7] is a harbinger of many other such studies that will follow in the wake of the Genome10K Project (http://www.genome10k.org), which sets out to assemble the genomic DNA sequences of 10 000 vertebrate species and to thereby cover every vertebrate genus. When overlaid on the well-documented diversity in reproductive strategies within the animal kingdom and the remarkable evolutionary divergence of reproductive anatomy and physiology that occurs in even the most closely related of species, these genomic data sets are in essence the virtual keys that may allow us to unlock many of the secrets that have long eluded reproductive biologists. This new age of comparative genomics is fast approaching, and it will have a huge impact on fields like reproductive biology, where sometimes dramatic biological adaptations have occurred in concert with rapid evolutionarily changes in the sequences of genes, like the Rhox family. It can be anticipated that these rapidly evolving genes will surface as the drivers or mediators of adaptive mechanisms that have ensured the diversity and, ultimately, the survival of species. Most important, these advances in knowledge are poised to occur quickly and can be expected to reveal new and unexpected ways of modifying the fertility of different species across the animal kingdom.
- Discussion
6
- 10.1111/mec.15801
- Feb 1, 2021
- Molecular Ecology
Species distributions are rapidly being altered by human globalisation and movement. As species are moved across biogeographic boundaries, human-mediated secondary contacts between historically allopatric taxa may promote hybridisation between closely related native and introduced species. The outcomes of hybridisation are diverse from strong reproductive barriers to gene flow to genome-wide admixture that may enhance (Fitzpatrick et al., 2010; Mesgaran et al., 2016; Valencia-Montoya et al., 2020) or impede (Kovach et al., 2016) invasive spread. For native species, introgressive hybridisation may disassemble locally adapted genomes, and in extreme cases, extensive asymmetric introgression may lead to the 'genomic extinction' of endemic diversity (Rhymer & Simberloff, 1996; Todesco et al., 2016). Undoubtedly, introgressive hybridisation can rapidly alter the evolution of introduced and endemic populations. This is a major conservation issue (Leitwein, Duranton, Rougemont, Gagnaire, & Bernatchez, 2020), with the greatest potential consequences on small, range-restricted native populations where introduced species may reach higher relative densities (Currat, Ruedi, Petit, & Excoffier, 2008). In this issue of Molecular Ecology, Blackwell et al. (2020) explore the history of divergence and admixture between the highly invasive Nile tilapia, Oreochromis niloticus, and a recently discovered Oreochromis lineage that is endemic to the coastal lakes of southern Tanzania. Oreochromis tilapias belong to the African cichlids and the most diverse family of vertebrates (Cichlidae), with almost 2000 species inhabiting the Great Lakes and river environments of Eastern Africa (Kocher, 2004; McGee et al., 2020). By analysing previously unrecognised cichlid diversity from southern lakes, Blackwell et al. (2020) provide novel evidence for how introgressive hybridisation with introduced species can alter native genetic makeup, illustrating the potential susceptibility of Tanzania's endemic biodiversity to genetic threats from introduced taxa.
- Research Article
154
- 10.1016/s0960-9822(02)01399-4
- Jan 1, 2003
- Current Biology
Hox Gene Loss during Dynamic Evolution of the Nematode Cluster
- Research Article
45
- 10.1016/j.parint.2007.09.007
- Oct 11, 2007
- Parasitology International
Hox genes and the parasitic flatworms: New opportunities, challenges and lessons from the free-living
- Peer Review Report
- 10.7554/elife.82290.sa0
- Oct 7, 2022
Full text Figures and data Side by side Abstract Editor's evaluation Introduction Results Discussion Materials and methods Data availability References Decision letter Author response Article and author information Metrics Abstract Functionally indispensable genes are likely to be retained and otherwise to be lost during evolution. This evolutionary fate of a gene can also be affected by factors independent of gene dispensability, including the mutability of genomic positions, but such features have not been examined well. To uncover the genomic features associated with gene loss, we investigated the characteristics of genomic regions where genes have been independently lost in multiple lineages. With a comprehensive scan of gene phylogenies of vertebrates with a careful inspection of evolutionary gene losses, we identified 813 human genes whose orthologs were lost in multiple mammalian lineages: designated 'elusive genes.' These elusive genes were located in genomic regions with rapid nucleotide substitution, high GC content, and high gene density. A comparison of the orthologous regions of such elusive genes across vertebrates revealed that these features had been established before the radiation of the extant vertebrates approximately 500 million years ago. The association of human elusive genes with transcriptomic and epigenomic characteristics illuminated that the genomic regions containing such genes were subject to repressive transcriptional regulation. Thus, the heterogeneous genomic features driving gene fates toward loss have been in place and may sometimes have relaxed the functional indispensability of such genes. This study sheds light on the complex interplay between gene function and local genomic properties in shaping gene evolution that has persisted since the vertebrate ancestor. Editor's evaluation The study provides a fundamental understanding of the driving forces behind gene losses in genome evolution and connects the propensity for gene losses to local genomic features like mutation rate and expression pattern. The methodology is compelling, as it identifies "elusive human genes" through independent gene losses in at least two mammalian lineages. The comparative genomics and statistical analyses are thorough and rigorous, making this study appealing to readers interested in exploring the global patterns and underlying mechanisms of gene fate evolution across the phylogenetic tree. https://doi.org/10.7554/eLife.82290.sa0 Decision letter Reviews on Sciety eLife's review process Introduction In the course of evolution, genomes continue to retain most genes with occasional duplications, while losing some genes (Blomme et al., 2006; Fernández and Gabaldón, 2020; Shen et al., 2018). This retention and loss can be interpreted as gene fate; genes are stably retained in the genome, but some factors may cause them to transition to a state where deletion occurs. Accordingly, identification of the factors allowing gene loss may facilitate our understanding of gene fate. Gene retention or loss has generally been considered to depend largely on the functional importance of the particular gene from the perspective of molecular evolutionary biology (Albalat and Cañestro, 2016; Bartha et al., 2018; Blanc et al., 2012; Liu et al., 2015; Olson, 1999; Sharma et al., 2018; Shen et al., 2018). Genes with indispensable functions have usually been retained with highly conserved sequences in genomes through rapid elimination of alleles that impair gene functions (Hirsh and Fraser, 2001; Krylov et al., 2003; Miyata et al., 1980; Pál et al., 2006). On the contrary, genes with less important functions are likely to accept more mutations and structural variations, which can degrade the original functions, leading to gene loss through pseudogenization or genomic deletion (Jordan et al., 2002; Yang et al., 2003). To date, gene loss has been imputed to the relaxation of functional constraints of individual genes. Gene loss has further been revealed to drive phenotypic adaptation in various organisms (Albalat and Cañestro, 2016; Olson, 1999), as well as in a gene knockout collection of yeasts in culture (Giaever and Nislow, 2014; Maclean et al., 2017). To uncover the association between fates and functional importance of the genes, molecular evolutionary analyses have been conducted at various scales, from gene-by-gene to genome-wide. A number of studies have revealed that the genes with reduced non-synonymous substitution rates (or KA values) and ratios of non-synonymous to synonymous substitution rates (KA/KS ratios) are less likely to be lost (Jordan et al., 2002; Yang et al., 2003). A genome-wide comparison of duplicated genes in yeast revealed larger KA values for those lost in multiple lineages than those retained by all the species investigated (Byrne and Wolfe, 2007). Other comprehensive studies of gene loss across metazoans and teleosts revealed that the genes expressed in the central nervous system are less prone to loss (Fernández and Gabaldón, 2020; Roux et al., 2017). These observations again suggest that gene fate depends on the functional constraints of a particular gene. Besides functional constraints, several studies have identified the genes lost independently in multiple lineages, revealing that the genomic regions containing these genes 'prefer' particular characteristics associated with structural instability (Cortez et al., 2014; Hughes et al., 2012; Lewin et al., 2021; Maeso et al., 2016). In mammals, tandemly arrayed homeobox genes derived from the Crx gene family were lost in multiple species (Lewin et al., 2021; Maeso et al., 2016). The findings suggest that genomic features containing tandem duplications facilitate unequal crossing over, leading to frequent gene loss. Mammalian chromosome Y, which contains abundant repetitive elements and continues to reduce in size, has lost a considerable number of genes (Cortez et al., 2014; Hughes et al., 2012). In the stickleback genome, a Pitx1 enhancer was independently lost in multiple lineages inhabiting freshwater due to its genomic location in a structurally fragile site, leading to recurrent loss of pelvic fins (Xie et al., 2019). Genes and genomic elements in such particular regions may be prone to loss in a more neutral manner than the relaxation of functional importance or via functional adaptations. Accordingly, these studies focusing on the particular genomic regions led us to search for the common features in genomes that potentially facilitate gene loss. Genome-wide scans have revealed heterogeneous distributions of a variety of sequence and structural features so far, for example, base composition (Bernardi and Bernardi, 1986; Cohen et al., 2005; Katzman et al., 2011), the frequency of repetitive elements (Korenberg and Rykowski, 1988; Medstrand et al., 2002), and DNA-damage sensitivity induced by replication inhibitors (Debatisse et al., 2012; Helmrich et al., 2006). However, the extent to which these characteristics are associated with gene fates has not been understood well at a genome-wide level. The accumulation of near-complete genome assemblies for various organisms facilitates comprehensive taxon-wide analysis of gene loss (Fernández and Gabaldón, 2020; Guijarro-Clarke et al., 2020; Rice and McLysaght, 2017). Along with this motivation, we recently performed a comprehensive analysis on the fate of paralogs generated via the two-round whole-genome duplications in early vertebrates (Hara et al., 2018a). The results revealed that the genes retained by reptiles but lost in mammals and Aves rapidly accumulated not only non-synonymous but also synonymous substitutions in comparison with the counterparts retained by almost all the vertebrates examined, indicating that those genes prone to loss show increasing mutation rates. Furthermore, these loss-prone genes were located in genomic regions with high GC contents, high gene densities, and high repetitive element frequencies. These findings suggest that the fates of those genes are influenced not only by functional constraints but also by intrinsic genomic characteristics. Because the findings were restricted to a set of particular genes, they prompted us to examine whether this trend is associated with gene fates on a genome-wide scale. In this study, we inferred molecular phylogenies of vertebrate orthologs to systematically search for the genes harboring different fates in the human genome. We previously referred to the nature of genes prone to loss as 'elusive' (Hara et al., 2018a; Hara et al., 2018b). In this study, we define the elusive genes as those that are retained by modern humans but have been lost independently in multiple mammalian lineages. As a comparison of the elusive genes, we retrieved the genes that were retained by almost all of the mammalian species examined and defined them as 'non-elusive,' representing those persistent in the genomes. We conducted a careful search for gene loss to reduce the false discovery rate (FDR), which is usually caused by incomplete sequence information (Botero-Castro et al., 2017; Deutekom et al., 2019). By comparing the genomic regions containing these genes, we uncovered genomic characteristics relevant to gene loss. We associated the elusive genes with a variety of findings from deep sequencing analyses of the human genome, including transcriptomics, epigenomics, and genetic variations. These data assisted us to understand how intrinsic genomic features may affect gene fate, leading to gene loss by decreasing the expression level and eventually relaxing the functional importance of 'elusive' genes. Results Identification of human 'elusive' genes We defined an 'elusive' gene as a human protein-coding gene that existed in the common mammalian ancestors but was lost independently in multiple mammalian lineages (Figure 1; see 'Materials and methods' for details). We searched for such genes by reconstructing phylogenetic trees of vertebrate orthologs and detecting gene loss events within the individual trees. To search for elusive genes, we paid close attention to distinguishing true evolutionary gene loss from falsely inferred gene loss caused by insufficient genome assembly, gene prediction, and orthologous clustering (Botero-Castro et al., 2017; Deutekom et al., 2019), as described below. Figure 1 Download asset Open asset Detection of 'elusive' genes. (a) Pipeline of ortholog group clustering and gene loss detection. (b) Definition of an elusive gene schematized with ortholog presence/absence pattern referring to a taxonomic hierarchy. Red and orange crosses denote the gene loss in the common ancestor of a taxon and the loss specific to a single species, respectively. (c) A representative phylogeny of the elusive gene encoding Chitinase 3-like 2 (CHI3L2). Taxa shown in the tree were used to investigate the presence or absence of orthologs. The Sciuromorpha, Hystricognathi, Eulipotyphla, Carnivora, and Chiroptera are absent from the tree, indicating that the CHI3L2 orthologs were lost somewhere along the branches framed in gray in the tree. In addition, the orthologs of many members of the Myomorpha were not found, suggesting that gene loss occurred in this lineage. We first produced highly complete orthologous groups comprised of nearly complete gene sets. We merged multiple gene annotations of a single species followed by assessments of the completeness of the gene sets (Figure 1a). Using these gene sets, we then created two sets of ortholog groups with different methods and merged them into a single set (Figure 1a). In searching for gene loss events, we restricted our study to those that occurred in the common ancestors of particular taxonomic groups. This procedure relieved false identifications of gene loss in a species or an ancestor of a lower taxonomic hierarchy caused by incomplete genomic information (Figure 1b). We integrated gene annotations from Ensembl, RefSeq, and the sequence repositories of individual genome sequencing projects to produce gene annotations for 114 mammalian and 132 non-mammalian vertebrates. From these, we selected the annotations of 101 and 90 species, respectively, that exhibited high completeness in the BUSCO assessment (Simão et al., 2015; Supplementary Table S1 in Supplementary file 1a). Using these gene sets, clustering of ortholog groups was conducted by OrthoFinder, and these groups were integrated into the ortholog groups provided by the Ensembl Gene Tree. This integration resulted in 50,768 vertebrate ortholog groups. Phylogenetic tree inference of the integrated ortholog groups and pruning of the individual trees based on gene duplications resulted in 17,495 mammalian ortholog groups that contained human genes. We classified the mammalian species into 15 taxonomic groups ranging from order to family (listed in Table S1; Supplementary file 1a). For the individual mammalian orthologs, we searched for the taxa in which the gene was absent in all the species examined (Figure 1b). We interpreted this gene absence as an evolutionary loss that occurred in the common ancestor of the taxon. Validating the gene loss through an ortholog search in genome assemblies and synteny-based ortholog annotations, we extracted the ortholog groups that were retained by humans but were lost independently in the common ancestors of at least two taxa (Figure 1c). Hereafter we call the human genes belonging to these ortholog groups 'elusive genes.' To compare these, we also selected the ortholog groups that contained all of the mammals examined including single-copy human genes. We called these 'non-elusive genes.' This comprehensive scan of gene phylogenies resulted in 813 elusive and 8050 non-elusive genes (Supplementary Table S2; Supplementary file 2). Genomic signatures of the human elusive genes The loss-prone nature of the elusive genes suggests a relaxation of their functional constraints. To uncover the molecular evolutionary characteristics associated with each elusive gene, we computed synonymous and non-synonymous substitution rates in coding regions, namely KS and KA, respectively, between human and chimpanzee and mouse orthologs for the elusive and non-elusive genes. In addition, we computed nucleotide substitution rates for introns (KI) between human and chimpanzee (Pan troglodytes) orthologs and compared them between the elusive and non-elusive genes. The results showed larger KA values in the ortholog pairs of the elusive genes than in those of the non-elusive genes (Figure 2a, Figure 2—figure supplement 1). This indicates a rapid accumulation of amino acid substitutions in the elusive genes, potentially accompanied by the relaxation of functional constraints. Our analysis further illuminated larger KS and KI values for the elusive genes than in the non-elusive genes (Figure 2b and c, Figure 2—figure supplement 1). Importantly, the higher rate of synonymous and intronic nucleotide substitutions, which may not affect changes in amino acid residues, indicates that the elusive genes are also susceptible to genomic characteristics independent of selective constraints on gene functions. Figure 2 with 1 supplement see all Download asset Open asset Genomic and evolutionary characteristics of elusive genes. Distributions of non-synonymous, synonymous, and intronic nucleotide substitution rates, namely KA (a), KS (b), and KI (c) values, respectively, between the human–chimpanzee orthologs of the elusive and non-elusive genes. Distribution of gene length (d) and GC content (e) of the human elusive and non-elusive genes. (f) Distribution of gene density in the genomic regions where the human elusive and non-elusive genes are located. The plots consist of 249 elusive and 5145 non-elusive genes that retained chimpanzee orthologs (a, b), 473 and 4626 of those which harbored introns aligned with the chimpanzee genome (c; see 'Materials and methods'), and all of the 813 elusive and 8050 non-elusive genes (d–f). Diamonds and bars within violin plots indicate the median and range from the 25th to 75th percentile, respectively. To further scrutinize the characteristics reflecting the genomic environment rather than gene function, we analyzed genomic characteristics that may distinguish the elusive from non-elusive genes. A comparison between these two categories revealed shorter gene-body lengths and higher GC contents of elusive rather than non-elusive genes (Figure 2d and e). Furthermore, a scan of intragenomic gene distribution revealed that the elusive genes were located in the genomic regions with high gene density compared with the non-elusive genes (Figure 2f). Our findings indicate that such elusive genes have distinct characteristics in the human genome. These genomic characteristics, as well as high nucleotide substitution rates, were consistent with the findings in our genome analyses using the amniote and elasmobranch genomes (Hara et al., 2018a; Hara et al., 2018b). Tracing elusiveness back along the vertebrate evolutionary tree The origins of the human elusive genes can be traced back along the evolutionary tree, at least to the mammalian common ancestor. To investigate possible antiquities of the genomic properties associated with elusive genes, we investigated their orthologs in non-mammalian vertebrates by scrutinizing the ortholog groups used for elusive gene identification. We found that 152 out of 813 elusive genes originated in mammalian lineages, and this proportion was larger than those of the elusive genes (65 out of 8050, p=2.50 × 10-110), indicating that the elusive genes are more abundant in recently born genes than non-elusive genes. We then selected 517 elusive and 7900 non-elusive genes that originated in the common ancestors of jawed vertebrates or earlier. These subsets allowed us to examine the degree of retention of non-mammalian vertebrate orthologs in the elusive and non-elusive genes. On average, approximately 40% of these elusive genes were found to be retained by non-mammalian vertebrates, while this proportion increased up to 90% for the non-elusive genes. (Figure 3—figure supplement 1a). In the coelacanth, gar, and shark, the orthologs of the elusive genes were less frequently retained by all the species than those of the non-elusive ones (Figure 3—figure supplement 1b). The results suggest that the origins of the loss-prone propensity of the elusive genes potentially date back to the period long before the emergence of the Mammalia. We further examined the genomic characteristics associated with the human elusive genes in the vertebrate orthologs. In all the species examined, orthologs of the elusive genes exhibited high GC content and compact gene bodies. Additionally, in most of these species, the orthologs of elusive genes were located in genomic regions with high gene density compared with orthologs of the non-elusive genes (Figure 3, Figure 3—figure supplement 2). In addition, we computed KS and KA values between the orthologs of the vertebrate species and their close relatives for elusive and non-elusive genes. In any of the species pairs except for avians, the orthologs of the elusive genes were found to harbor higher KA and KS values than those of the non-elusive gene orthologs (Figure 3, Figure 2—figure supplement 1). These observations indicate that these genomic characteristics probably originated before the emergence of gnathostomes, a monophyletic group of chondrichthyan and bony vertebrates, and have been retained for approximately 500 million years. Figure 3 with 2 supplements see all Download asset Open asset Long-standing characteristics of elusive genes. Retention of the genomic and evolutionary characteristics of the human elusive genes across vertebrates. The individual round squares with arrows indicate significant increases or decreases of the distribution of particular characteristics in the orthologs of the human elusive genes and their flanking regions compared with those of the non-elusive genes in these selected vertebrate genomes. For the chimpanzee and mouse genomes, KA and KS values were computed between the human elusive genes and the orthologs of these mammals. For non-mammalian species, these values were computed with ortholog pairs for the elusive/non-elusive genes between the corresponding species and their closely related species: turkey for chicken, green anole for central bearded dragon, and whale shark for bamboo shark. Distributions of these metrics for non-human species are shown in Figure 2—figure supplement 1 and Figure 3—figure supplement 2. Species name: mouse, Mus musculus; chicken, Gallus gallus; central bearded dragon, Pogona vitticeps; Western clawed frog, Xenopus tropicalis; coelacanth, Latimeria chalumnae; spotted gar, Lepisosteus oculatus; bamboo shark, Chiloscyllium plagiosum. Abundant polymorphism in elusive genes The observation of large KS and KA values in the elusive genes prompted us to examine the extent to which these genes have accommodated genetic variations in modern humans. Large-scale human genome resequencing projects have identified a huge number of genetic variations, from rare to common, and from single-nucleotide variants (SNVs) to chromosome-scale structural variants, facilitating tackling this issue. We retrieved copy number variants (CNVs) and rare SNVs in the human genome from the Database of Genomic Variants, release 2016-08-31 (MacDonald et al., 2014) and dbSNP release 147 (Sherry et al., 2001), respectively, and computed their densities in the individual genic regions. We found that the genic regions of the human elusive genes contained abundant rare SNVs, as well as deletion and duplication CNVs, compared with those of the non-elusive genes (Figure 4a–c). This result suggests that genomic regions containing the elusive genes are not only prone to loss but also to duplication. Figure 4 Download asset Open asset Genetic variations of the elusive and non-elusive genes within human populations. Comparison of the density of rare single-nucleotide variants (SNVs) (a), deletion copy number variants (CNVs) (b), duplication CNVs (c), and Z-scores of synonymous (d), missense (e), and loss-of-function variants (f). We used opposite numbers of the Z-scores in d–f so that the elusive genes have higher values than non-elusive genes as in Figure 2a, b, c, e, f and Figure 3a–c. (a–c) 813 elusive genes and 8050 non-elusive genes were used. (d–f) 544 elusive genes and 7303 non-elusive genes for which genetic variants were available in GnomAD were used. Diamonds and bars within violin plots indicate the median and range from 25th to 75th percentile, respectively. To evaluate the functional consequences of abundant genetic variants in the elusive genes, we investigated genetic variations stored in the gnomAD v. 2.1 database, a repository containing >120,000 exome and >15,000 whole-genome sequences of human individuals (Karczewski et al., 2021). This database classifies SNVs in coding regions into three and the loss-of-function contains and mutations in The gnomAD a an representing the of SNVs for individual and values denote or more mutations in a coding than (Figure Accordingly, the for mutations and loss-of-function mutations of the individual genes indicates the degree of larger values genes to while ones suggest functional We found lower Z-scores of missense and loss-of-function mutations opposite numbers of Z-scores in Figure and in the human elusive genes than in the non-elusive genes, suggesting that the elusive genes are more and potentially to Additionally, opposite numbers of Z-scores of synonymous mutations of the human elusive genes were higher than those of the non-elusive genes (Figure This the high mutability of genomic regions containing elusive genes, as in the KS of elusive genes To further investigate how the human elusive genes have functional we examined their expression For this we compared gene expression of the from the database v. et al., between the elusive and non-elusive genes. For individual genes, we computed the million values these as the expression level. For expression we which is as an of species in the based on the proportion of values across the As shown in the density plots of the individual genes these two in Figure most of the non-elusive genes large and Thus, most non-elusive genes are expressed at By the density of the elusive genes an with and values, indicating that the genes in this were not at least in The also showed of values, which contained the genes expressed in a single or a A analysis was performed with the single data et al., revealing that the expression of the elusive and non-elusive genes for the were with those of the (Figure Our findings that some elusive genes harbor and restricted expression that less which are in the non-elusive genes. Figure with 1 supplement see all Download asset Open asset of elusive and non-elusive genes. The density plots of the expression and of elusive and non-elusive genes. The numbers of the elusive/non-elusive genes and those for which the expression were available are in each were computed via 2 × 2 numbers of elusive and non-elusive genes with 1 and The median million of each of the across individuals was retrieved from the database et al., and values of the were retrieved from the database et al., For the individual genes, and values were computed using these nature of elusive genes Our of the and restricted expression patterns of elusive genes prompted us to properties in this transcriptional regulation. we retrieved data on a variety of human from a genome including a repository that the comprehensive annotations of functional elements in the human genome 2012). Using this we the features of the genomic regions containing elusive genes (Figure Figure with supplements see all Download asset Open asset features of the elusive genes. Comparison of the distribution of density (a), length of the including the elusive or non-elusive genes (b), the replication based on (c), and with the computed from of the analyses were performed by using the sequencing data available Supplementary file 1b). (d) and were performed with was performed with and was performed with In the elusive gene indicates the elusive genes with restricted 1; Figure for individual indicate the comparison between the elusive and non-elusive genes and the between the elusive genes with 1 and those with 1 The results for are shown in Figure supplements For the individual characteristics, for multiple was performed for comparison in each We compared densities based on the for using sequencing an of regions in the genome, in gene and flanking regions between the elusive and non-elusive genes. In all of the examined in the results showed in the genomic regions including the elusive genes than in those including non-elusive genes, indicating that the elusive genes are likely to in genomic regions (Figure Figure supplement 1). We also searched for genomic elements with frequent potentially as et al., 2014) that the elusive or non-elusive genes. The result showed that a higher of the elusive genes of the than the non-elusive genes for all the investigated (Figure Figure supplement