Comparison of protein-coding genes and pseudogenes among different chicken breeds.

a. Venn diagram of the protein-coding genes of our four indigenous chickens. b. Venn diagram of the protein-coding genes of our four indigenous chickens and RJF. c. Comparison of the protein-coding genes among each pair of our four indigenous chickens and RJF. d. Venn diagram of the pseudogenes of our four indigenous chickens. e. Venn diagram of the pseudogenes of our four indigenous chickens and RJF. f. Comparison of the pseudogenes among each pair of our four indigenous chickens and RJF. g. Venn diagram of the protein-coding genes of the four HiFi-reads assembled genomes. h. Venn diagram of the pseudogenes of the four HiFi-reads assembled genomes.

Summary of annotated pseudogenes in the four indigenous chicken genomes in comparison with those in GRCg6a

Summary of annotated protein-coding genes, pseudogenes and substitution rates in the four PacBio HiFi-reads assembled chicken genomes

Pseudogenization mutations tend to occur at the two ends of CDSs.

a-b. Probability of the most upstream pseudogenization mutations (red line) in 100 evenly divided CDS segments from the 5’-ends to the 3’-ends of the parental genes of the pseudogenes, mean rates of synonymous substitution in 100 evenly divided CDS segments from the 5’-ends to the 3’-ends of the true genes (blue line) and pseudogenes (purple), and mean identity of the true genes in 100 evenly divided CDS segments from the 5’-ends to the 3’-ends of the genes (green line). (a: For our four indigenous chickens; b: For the four HiFi-reads assembled genomes). c-d. Start position of the “CDS” of the pseudogenes with respect to the nucleotide positions of their parental genes starting with 0 with the downstream positions being positive integers (c: For our four indigenous chickens; d: For the four HiFi-reads assembled genomes). e-f. End positions of the “CDS” of the pseudogenes with respect to the nucleotide positions of their parental genes ending with 0 with the upstream positions being negative integers (e: For our four indigenous chickens; f: For the four HiFi-reads assembled genomes). g-h. Violin plots of the dN/dS values of all true genes, all pseudogenes, pseudogenes with the most upstream pseudogenization occurring in the first 10%, the intermediate 80% and the last 10% of the CDSs (g: For our four indigenous chickens; h: For the four HiFi-reads assembled genomes). i-j. Number of predicted miRNA binding sites per 100pb along the CDSs and 3’ UTRs of the true genes (red line) and pseudogenes (green line). In the figure, ‘0’ represents the end positions of the CDSs, the positive numbers represent the relative positions of 1,000 bp sequences downstream of the end of CDSs, and the negative numbers represent the relative positions of the CDSs with respect to the ends of CDSs (i: For our four indigenous chickens; j: For the four HiFi-reads assembled genomes).

Spectrum analysis of pseudogenes in our four chicken breeds and the four HiFi-reads assemblies.

a-d. Number of pseudogenes with the indicated fixation rate of the most upstream pseudogenization mutations in the chicken populations. e-h. Number of pseudogenes in each of our indigenous chicken genomes with the indicated identity with their parental genes. i-l. Number of pseudogenes in each of the four HiFi-reads assembled genomes with the indicated identity with their parental genes. m-n. Distribution of divergence times of pseudogenes in each chicken breed since the divergence from their last common ancestral genes shared with the RJF (GRCg6a).

Mean substitution rates (number of fixations per year per site) of protein-coding genes and pseudogenes in the four chicken breeds with respect to orthologous sequences in RJF (GRCg6a)

Analysis of frequency spectrums of SNPs.

a. Principal component analysis of the RJF subspecies, our indigenous chickens and the RJF individual (GRCg6a) based on their SNP profiles. b-d. Genetic structures of the RJF subspecies, our indigenous chickens and the RJF individual (GRCg6a) estimated using the ADMIXTURE program for K=2, 3, …, 15.

Evolutionary pattern of our four indigenous chickens and the RJF.

a. A hypothetical scenario for loss-of-functions of the five chickens since their divergence from MRCA A1 and A2. b-e. Number of missing genes in each of our indigenous chicken genomes with the indicated reads coverage on their functional versions. The dashed red lines are the cumulative density function (CDFs) of coverage ratios. f-i. Number of missing genes with the indicated missing rate in the chicken populations.

Distribution of pseudogenes and missing genes on each chromosome of the four indigenous chicken genomes.

a, b. Number of pseudogenes (a) and missing genes (b) on each chromosome of the chicken genomes. c, d. Density of pseudogenes (c) and missing genes (d) on each chromosome of the chicken genomes. e, f. Ratio of the number of pseudogenes over the number of genes (e), and ratio of the number of missing genes over the number of genes (f), on each chromosome of the chicken genomes. g. Comparison of G/C contents of true genes, pseudogenes and missing genes in the chicken genomes. Statistical tests were done using one-tailed t-test.

Evolutionary relationships of each chicken breed.

a. Heatmap of two-way hierarchical clustering of the dispensable genes in MRCA A1 that are either completely lost or pseudogenized in at least one of our four indigenous chickens and RJF based on their appearance as an intact form (1, brown), absence (0, white) and as a pseudogenized form (-1, blue) in the five chicken genomes. b. Heatmap of two-way hierarchical clustering of the dispensable genes in MRCA A1 that are either completely lost or pseudogenized in at least one of the four HiFi-reads assembled genomes and RJF based on their appearance as an intact form (1, brown), absence (0, white) and as a pseudogenized form (-1, blue) in the five chicken genomes. c. Neighbor-joining phylogenetic tree of our four indigenous chickens and RJF, constructed using the occurring patterns of the dispensable genes in their genomes. The numbers on the branches are Euclid distance between the pattern vectors. d. Neighbor-joining phylogenetic tree of our four indigenous chickens and RJF, constructed using the 6,744 essential protein-coding genes in their genomes and the quail genome. The numbers on the nodes are bootstrapping value for 1,000 repeats. e. Neighbor-joining phylogenetic tree of the four HiFi-reads assembled genomes and RJF, constructed using the occurring patterns of the dispensable genes in their genomes. The numbers on the branches are Euclid distance between the pattern vectors. f. Neighbor-joining phylogenetic tree of the four HiFi-reads assembled genomes and RJF, constructed using the 6,323 essential protein-coding genes in their genomes and the quail genome. The numbers on the nodes are bootstrapping value for 1,000 repeats.