Histone variants are paralogs that replace canonical histones in nucleosomes, often imparting novel functions. However, how histone variants arise and evolve is poorly understood. Reconstruction of histone protein evolution is challenging due to large differences in evolutionary rates across gene lineages and sites. Here we used intron position data from 108 nematode genomes in combination with amino acid sequence data to find disparate evolutionary histories of the three H2A variants found in Caenorhabditis elegans: the ancient H2A.ZHTZ-1, the sperm-specific HTAS-1, and HIS-35, which differs from the canonical S-phase H2A by a single glycine-to-alanine C-terminal change. Although the H2A.ZHTZ-1 protein sequence is highly conserved, its gene exhibits recurrent intron gain and loss. This pattern suggests that specific intron sequences or positions may not be important to H2A.Z functionality. For HTAS-1 and HIS-35, we find variant-specific intron positions that are conserved across species. Patterns of intron position conservation indicate that the sperm-specific variant HTAS-1 arose more recently in the ancestor of a subset of Caenorhabditis species, while HIS-35 arose in the ancestor of Caenorhabditis and its sister group, including the genus Diploscapter. HIS-35 exhibits gene retention in some descendent lineages but gene loss in others, suggesting that histone variant use or functionality can be highly flexible. Surprisingly, we find the single amino acid differentiating HIS-35 from core H2A is ancestral and common across canonical Caenorhabditis H2A sequences. Thus, we speculate that the role of HIS-35 lies not in encoding a functionally distinct protein, but instead in enabling H2A expression across the cell cycle or in distinct tissues. This work illustrates how genes encoding such partially-redundant functions may be advantageous yet relatively replaceable over evolutionary timescales, consistent with the patchwork pattern of retention and loss of both genes. Our study shows the utility of intron positions for reconstructing evolutionary histories of gene families, particularly those undergoing idiosyncratic sequence evolution.
Read full abstract