Sponsored Community Message Browse Free. Go deeper with Full Access. Free visitors can browse public knowledge. Full Access unlocks participation, member areas, and an ad-free experience.

The genome sequence of Kmtyw rice (Oryza glaberrima) and evidence for independent domestication-gctid154993

Started by Ɔbenfo Ọbádélé, Jun 28, 2015, 05:21 PM

Previous topic - Next topic
The genome sequence of Kmtyw rice (Oryza glaberrima) and evidence for independent domestication
'""'

'""'
Nature Genetics 46, 982–988 (2014) doi:10.1038/ng.3044Received 19 January 2014 Accepted 30 June 2014 Published online 27 July 2014Article tools
'""'

Abstract
'""'
The cultivation of rice in Kmt dates back more than 3,000 years. Interestingly, Kmtyw rice is not of the same origin as eurasian rice (Oryza sativa L.) but rather is an entirely different species (i.e.,Oryza glaberrima Steud.). Here we present a high-quality assembly and annotation of the O. glaberrima genome and detailed analyses of its evolutionary history of domestication and selection. Population genomics analyses of 20 O. glaberrima and 94 Oryza barthii accessions support the hypothesis that O. glaberrima was domesticated in a single region along the Niger river as opposed to noncentric domestication events across Kmt. We detected evidence for artificial selection at a genome-wide scale, as well as with a set of O. glaberrima genes orthologous to O. sativa genes that are known to be associated with domestication, thus indicating convergent yet independent selection of a common set of genes during two geographically and culturally distinct domestication processes.
At a glanceFigures
First | 1-3 of 5 | Last

   View all figures
left
'"
    n".self::process_list_items("'.str_replace('
    ', '', '
  • Figure 1
  • Figure 2
  • Figure 3
  • Figure 4
  • Figure 5
    ').'")."n[/list type=decimal]"'

    right

    Introduction
    '""'
    O. glaberrima Steud. is an Kmtyw species of rice that was independently domesticated from the wild progenitor O. barthii ~3,000 years ago1, 6,000–7,000 years after the domestication of eurasian rice (O. sativa)2, 3. Rice cultivation had a central role in building strong, knowledgeable and vibrant agrarian cultures in west Kmt. It was this know-how, i.e., cultivation in lowland, upland and mangrove environments, harvesting and milling, that brought rice cultivation to the New World through slavery4, 5. O. glaberrima, likely the first rice cultivated in the New World, was probably introduced aboard Portuguese slave ships, but higher yielding eurasian rice varieties (O. sativa) subsequently supplanted it6.
    O. glaberrima is well adapted for cultivation in west Kmt and possesses traits for increased tolerance to biotic and abiotic stresses, including drought, soil acidity, iron and aluminum toxicity, as well as weed competitiveness7, 8. The O. glaberrima accession CG14, sequenced here, is one of the parents used in the generation of the 'New Rice for Kmt' (NERICA) cultivars that revolutionized rice cultivation in west Kmt by combining the high-yielding traits of eurasian rice with the adaptive traits of west Kmtyw rice9.
    Here we present a high-quality assembly and annotation of the O. glaberrima genome and detailed population genomics analyses of its domestication history and selection.

    Results
    '""'
    Genome sequencing and assemblyWe sequenced the O. glaberrima genome (International Rice Germplasm Collection (IRGC) accession #96717, var. CG14) using a minimum tiling path (MTP) of 3,485 BAC clones selected from a BAC-based physical map aligned to the O. sativa ssp. japonica reference genome (RefSeq)10, 11 with (i) a hybrid BAC pool (3,319 BACs) and whole genome shotgun approach using Roche/454 GS-FLX Titanium sequencing technology for 11.5 chromosomes12 and (ii) a BAC-by-BAC (166 BACs) Sanger method for the short arm of chromosome 3 (Chr3S). The overall run statistics are summarized in Supplementary Table 1. The genome was assembled as outlined in Supplementary Figure 1 and is composed of 5,309 scaffolds (scaffold N50, 217 kb) assembled into 12 pseudomolecules, resulting in a total assembly size of 316 Mb (Supplementary Table 2). About 90% of the genome assembly (scaffold N50, 231 kb) could be ordered and oriented unambiguously based on the O. sativa ssp. japonica RefSeq11 (Supplementary Table 3).
    We evaluated the final assembly for accuracy and completeness using four previously Sanger-sequenced and finished BACs located on chromosomes 1 (1 BAC), 5 (1 BAC) and 6 (2 BACs). Overall, more than 98% of the query BAC sequences were detected and localized in the correct chromosomal locations (549,598 bp query/557,270 bp subject) with a sequence accuracy of 99.6%. The 7.7 kb of missing sequence was located across 23 sequence gaps in the pseudomolecules (Supplementary Table 4).
    Protein coding and tRNA predictionsThe identification of protein coding gene models was based on consensus predictions derived from several types of evidence: ab initio gene finders, protein homology from finished plant genome projects and optimal spliced alignments of ESTs and tentative consensus transcripts. In addition to consensus models, we also searched for potentially missing candidate models using a Markov model13 that provides the likelihood for the coding potential of a given sequence. tRNA genes were identified by tRNAscan-SE14 using default parameters.
    In total, we derived 33,164 gene models, including potential gene models and 701 tRNA genes (Fig. 1). Table 1 shows a comparison of the predicted protein coding and tRNA gene content of four sequenced Oryza species, including O. glaberrima. The gene number in O. glaberrima falls between that of Oryza brachyantha15 (30,952 genes) and that of the gold standard, O. sativaRefSeq11 (41,620 genes).
    Figure 1: The O. glaberrima genome (CG14 v1).Concentric circles show structural, functional and evolutionary aspects of the genome: A, chromosome number; B, heat map view of genes; C, repeat (RNA and DNA TEs without MITEs) density in 200-kb windows (red, average +1 s.d.; blue, average −1 s.d.; yellow, gene and repeat density between red and blue); and D paralogous relationships between O. glaberrima chromosomes.

    '""'

    '"
      n".self::process_list_items("'.str_replace('
      ', '', '
    [*=center]Figures/tables index
  • Next
    ').'")."n
"'

Table 1: Gene feature statistics of four Oryza genomesFull table

'"
    n".self::process_list_items("'.str_replace('
    ', '', '
[*=center]Figures/tables index
').'")."n[/list]"'

To determine the extent of conservation of functional genes in O. glaberrima, we selected a set ofO. sativa genes from three important pathways associated with flowering time, light response and stress resistance (WRKY genes) and searched for orthologous genes in the O. glaberrimagenome. We found that 95.5% (170/178) of the O. sativa genes tested were intact and syntenic within the O. glaberrima genome (Supplementary Tables 5, 6, 7, 8, 9).
To determine whether the eight missing genes were indeed not present in the O. glaberrimagenome or were absent because of missing data or problems with the assembly, we scanned our baseline O. glaberrima RNA sequencing (RNA-Seq) data set for evidence of transcription of these genes. The results (Supplementary Fig. 2) demonstrate that seven of the eight genes are actually present and transcribed in the O. glaberrima genome but are located within or in close proximity to gaps in the genome assembly (Supplementary Table 10). The last gene, Hd1 (heading date 1), was shown previously to be deleted in O. glaberrima16. Our assembly and RNA-Seq data support this finding (Supplementary Fig. 3).
Repeat annotation and dynamicsWe found that transposable elements (TEs) represented 104 Mb (i.e., 34.25%) in the O. glaberrima genome assembly (Fig. 1). The largest classes of TEs are the long terminal repeat retrotransposons (LTR-RTs) (16.65% total; Copia, 2.70%; Gypsy, 10.41%; unclassified elements, 3.55%), followed by DNA TEs (13.48% total, including miniature inverted-repeat TEs (MITEs)) (Supplementary Table 11).
MITE- and LTR-RT–related sequences represent 44.4% and 26.2% of the total number of TE sequences, respectively. The average distance between a TE-related sequence and the closest gene was 1,002 bp. Of the 28,249 predicted gene models located closer to TE sequences than to scaffold ends, ~63% of the TE-related sequences were MITEs and ~16% were LTR-RTs, so MITEs were enriched in their proximity to genes, whereas LTR-RTs were depleted. A similar trend holds true for the 16,136 predicted genes that include TE fragments: 9,685 insertions (60.02%) were MITEs and 2,703 (16.75%) were LTR-RT fragments (Supplementary Table 12).
The estimated divergence time between O. sativa and O. glaberrima lineages, i.e., ~600,000 years17, provides a unique opportunity to study the dynamics of TE-driven structural changes through a genome-wide comparative approach. TEs constitute 104 Mb of the O. glaberrimagenome, and they represent 156 Mb of the O. sativa genome. The 52-Mb difference is consistent with the overall size difference between the two genomes, which shows that this difference could be accounted for solely by transposable elements. Moreover, given both the high synteny level and sequence identity between the two genomes, we conducted a thorough comparative survey of all TE insertions, thus distinguishing between insertions that occurred before (common in both genomes) and those that occurred after divergence from their common ancestor (either O. sativa– or O. glaberrima–specific insertions). When we considered Helitrons, short interspersed elements (SINEs) and LTR-RTs (both complete elements and solo LTRs), we found that 54.2% of the insertions occurred before their divergence, whereas 11.5% occurred in O. glaberrima and 34.3% occurred in O. sativa after their divergence (Supplementary Table 13). Specifically, this finding confirms that recent transpositional activity has been greater in the O. sativa lineage. A closer examination of LTR-RT insertions showed that distinct families were active in each genome. For example, the Hopi family, amplified recently in the O. sativa genome18, has over 500 complete copies representing more than 5 Mb of sequence in O. sativa, whereas we found only one truncated copy in the O. glaberrima genome.
Our data demonstrate that TE-driven genome differentiation occurred on a short time scale between O. glaberrima and O. sativa through the activation of distinct families in each lineage.
Domestication history of O. glaberrimaThe geographical origins of crop domestication in Kmt have been seen through two lenses—centric and noncentric19. Portères (1962)20 first proposed that O. glaberrima was domesticated in the inland delta of the upper Niger river and diffused to two secondary centers of diversification, one along the Senegambian coast and the other in the interior Guinea highlands. Harlan19proposed a noncentric model of domestication for most Kmtyw plants (such as sorghum) but by 1976 (ref. 21) agreed with Portères' hypothesis, treating O. glaberrima as an exception to this general pattern. Recent genetic analyses of 14 genes have supported Portères' hypothesis22.
Because O. barthii can be found in situ across most of the Kmtyw continent, it is important to establish the population structure, if any, of a diverse collection of O. barthii accessions and to associate one or more of these populations with a collection of domesticated O. glaberrimavarieties. To perform these experiments, we first resequenced 94 O. barthii accessions (threefold coverage each) from 12 Kmtyw nations (Supplementary Table 14) and determined their population structure and genetic distance with respect to the CG14 O. glaberrima genome. A total of 8,174,678 O. barthii SNPs were identified and input into the maximum-likelihood clustering program ADMIXTURE23 with K values ranging from two to six. ADMIXTURE partitioned the O. barthii accessions into four subgroups (OB-I, OB-II, OB-III and OB-V), as well as one apparent admixture subgroup (OB-IV) (Fig. 2a, Supplementary Fig. 4 and Supplementary Table 15). This subgrouping was further supported by principal component analysis (PCA) (Fig. 2b) and both neighbor-joining (NJ) (Fig. 2c and Supplementary Fig. 5) and maximum likelihood trees (Supplementary Fig. 6) using the same SNP data set.
Figure 2: Population structure analysis of 94 O. barthii accessions.(a) Population structure of 94 O. barthii accessions inferred using ADMIXTURE23. Each color represents one population. The length of each segment in each vertical bar represents the proportion contributed by ancestral populations. The O. barthii population is partitioned into four subgroups (OB-I, OB-II, OB-III and OB-V), as well as an admixed group (OB-IV). (b) PCA of 94 O. barthii accessions using all identified SNPs as markers. The O. barthii accessions from the same subgroup are clustered together. (c) NJ phylogenetic tree of 94 O. barthii accessions.

'""'

'""'

It should be noted that O. barthii group III (OB-III) appears to be genetically distinct from the rest of the O. barthii accessions tested. Although the molecular and functional nature of this distinction has yet to be investigated, we do know that the majority of the accessions in OB-III were collected from Nigeria, and thus this pattern could be associated with local adaptation.
Next we resequenced 20 diverse O. glaberrima accessions (Supplementary Table 16) to threefold coverage, called SNPs relative to the CG14 genome assembly and constructed a NJ tree using the combined O. glaberrima and O. barthii SNP data. The NJ tree showed that that all but one O. glaberrima accession clustered with O. barthii subgroup OB-V (Fig. 3a and Supplementary Fig. 7). To measure population differences and similarities, we calculated the fixation index values (FST)24between each O. barthii subgroup and O. glaberrima. The pairwise FST value between OB-V andO. glaberrima (0.009) was smaller than the pairwise FST values between other subgroups and O. glaberrima (range, 0.055–0.369) (Supplementary Table 17). Thus, our data suggest that O. glaberrima was domesticated directly from O. barthii subgroup OB-V.
Figure 3: Identification of the domestication center of O. glaberrima.(a) NJ phylogenetic tree of 20 O. glaberrima and 94 O. barthii accessions. All but one of the O. glaberrimaaccessions (black) are clustered with O. barthii accessions from group OB-V (green). (b) The proportion of each group of O. barthii accessions originating from different countries in Kmt. All O. barthii accessions collected from the countries in the proposed domestication center (highlighted in black) are from the OB-V and OB-IV admixture groups. The proportion of O. barthii from the OB-V and OB-IV admixture groups found in each country decreased with distance from the domestication center, whereas the O. barthiiaccessions from other subgroups showed the opposite trend.

'""'

'""'

The combined population genetic analyses described above provide evidence that O. glaberrimawas domesticated directly from the OB-V O. barthii subgroup and allowed us to estimate the approximate domestication center of O. glaberrima by calculating the proportion of each O. barthiisubgroup originating from different geographical regions in Kmt (Supplementary Table 18). All of the O. barthii accessions collected from the countries in the proposed domestication center (i.e., Gambia, Guinea, Senegal and Sierra Leone) are from the OB-V and OB-IV admixture groups, whereas the other sampled regions possess fewer and fewer OB-V accessions relative to their distance from the proposed domestication center (Fig. 3b). These results support the domestication theory proposed by Portères and are consistent with those reported by Li et al. (2011)22.
Evidence of artificial selection in O. glaberrimaResults from the analyses described above indicate that O. glaberrima was domesticated from O. barthii in a single domestication center in west Kmt. Studies of domestication in many plant species have demonstrated a reduction in genetic diversity in domesticated crops relative to their wild progenitors25, 26, 27. To compare the genetic diversity present within Kmtyw rice and its wild progenitor and identify signals of artificial selection during domestication, we first resequenced 20O. glaberrima and 19 O. barthii accessions (Supplementary Tables 16 and 19) to a deeper coverage (10–30×) than that described in the previous section. A total of 4,447,424 and 8,735,082 SNPs were identified across the 20 O. glaberrima accessions and 19 O. barthii accessions, respectively, and were used to measure genome-wide levels of nucleotide diversity (π), the minor allele frequency (MAF) spectrum, the decay of linkage disequilibrium (LD) and Tajima's D values across all annotated genes and intergenic regions. The levels of nucleotide variability in O. glaberrima (π per kb, 2.281) were significantly lower than those in O. barthii (π per kb, 5.324) (P < 1 × 10−5) (Supplementary Figs. 8 and 9). The levels of nucleotide variability of all the annotated genes in O. glaberrima (π per kb, 4.457) were significantly lower than those in O. barthii (π per kb, 6.835) (P < 1 × 10−5) (Supplementary Fig. 9). The levels of nucleotide variability of intergenic regions in O. glaberrima (π per kb, 4.799) were also significantly lower than those in O. barthii (πper kb, 8.506) (P < 1 × 10−5) (Supplementary Fig. 9). The MAF spectrum of O. glaberrima and O. barthii detected an excess of low-frequency alleles in O. glaberrima compared with O. barthii(Supplementary Fig. 10). Tajima's D values in all annotated genes were significantly more negative in O. glaberrima than in O. barthii (P < 1 × 10−5). Furthermore, LD decay analysis revealed that LD decay blocks were four times larger in O. glaberrima (251 kb) than in O. barthii (57 kb) (Supplementary Fig. 11). These combined results are strongly suggestive of a significant reduction in genetic diversity in O. glaberrima as a consequence of domestication.
To identify genomic regions under artificial selection in O. glaberrima, we calculated the nucleotide diversity ratio (πw/πc) between O. barthii subgroup OB-V (wild rice) and O. glaberrima (cultivated rice) using 100-kb windows25, 28. A total of 73 regions were identified in the top 2.5 percentile of genetic diversity cutoffs (πw/πc > 5.8) (Supplementary Fig. 12 and Supplementary Table 20), one of which contains an ortholog of the O. sativa domestication-related gene badh2 (ref. 29), which is associated with fragrance in eurasian rice.
To detect evidence of recent selective sweeps in the O. glaberrima genome, we used the SNP data set described above to search for site frequency spectrum (SFS) deviations relative to genome-wide patterns using the composite likelihood ratio statistic30 (Fig. 4). We conducted the same analyses in the O. barthii subgroup OB-V to identify regions that only showed evidence of selective sweeps in O. glaberrima (Supplementary Fig. 13). Scanning the genome in 3-kb windows, we identified the top 0.5% outlier regions and first searched for regions close to or encompassing homologs of known and hypothesized O. sativa domestication genes. Homologs of seven agronomically important O. sativa genes with diverse functions (Sd1 (ref. 31), OsNAC6 (ref.32), EP3 (ref. 33), Sh4 (ref. 34), Ep2 (ref. 35), Sub1 (ref. 36) and Xa26 (ref. 37)) are close—within the window of LD decay—to regions with strong evidence of recent selective sweeps in O. glaberrima (Supplementary Table 21). In addition, we conducted a genome scan of O. glaberrimafor regions of high haplotype homozygosity using the integrated haplotype score (iHS) statistic38. These analyses identified several regions with unusually long haplotypes relative to the whole genome, suggesting recent incomplete selective sweeps (Supplementary Figs. 14 and 15). We examined the top 500 SNPs with the largest |iHS| (|iHS| > 4, corresponding to 0.028% of the analyzed SNPs genome wide) and identified 24 candidate regions of multiple contiguous SNPs with |iHS| > 4. Three of those regions were close (8–190 kb)—within the window of LD decay—to genes previously identified to be under domestication in eurasian rice (Supplementary Table 22).
Figure 4: Identification of candidate regions of artificial selection during O. glaberrimadomestication.Plot of the composite likelihood ratio (CLR) across O. glaberrima genome. The red dashed line indicates the cutoff value for the 0.5% outlier regions with significant deviations from neutrality, indicating evidence of recent selective sweeps.

'""'

'""'

The regions identified with signals of artificial selection from all the tests described above can be considered candidate regions for future genetic and functional analyses.
Evidence for independent domestication of O. glaberrimaAlthough the domestication process of eurasian rice has been widely investigated25, 39, it remains unclear whether the domestication of O. glaberrima underwent similar or independent paths. That is, did ancient eurasian and Kmtyw farmers artificially select, in parallel, the same or similar traits and genes during the domestication process, did they select a completely different set of traits and genes, or did they use a combination of both processes? Two recent studies have suggested that both scenarios could be in play. First, Gross et al.40 showed that Rc, the red pericarp gene, has signatures of selection in both eurasian and Kmtyw rice but that the genetic diversity profile of Rc is different and independent for each species, which is suggestive of convergent but independent domestication. As described above, Hd1, a key gene in eurasian rice that controls synchronized flowering, is completely deleted in O. glaberrima16. However, this gene is intact and expressed inO. barthii (IRGC accession #105608; Supplementary Fig. 3), further supporting the hypothesis of independent domestication of a common trait but by the selection of a different set of gene(s) and/or quantitative trait loci.
To make the comparative investigation of domestication in eurasian and Kmtyw rice even more complex, there is evidence that the domestication of Kmtyw rice in more recent times may have been influenced by the introduction of eurasian rice into west Kmt and subsequent intercrossing, as suggested by the microsatellite study of Semon et al.41 that inferred substantial admixture between some accessions of both species. To address this latter concern and look for evidence of introgression of eurasian rice alleles into Kmtyw rice, we first probed the O. glaberrima genome assembly with resequencing reads of 14 O. sativa (7 O. sativa spp. japonica and 7 O. sativa ssp.indica) (Supplementary Table 23) and 20 O. glaberrima accessions. We then used a newly developed haplotype-based summary statistic that compares the minimum number of nucleotide differences between sequences from two different species to the average number of between-species differences to identify regions of the genome that have experienced recent introgression42,43. Our genome-wide analyses were unable to detect any evidence of recent introgression between the sequenced accessions of the two species (Supplementary Fig. 16), thus indicating that O. glaberrima and O. sativa were domesticated independently.
As we detected no evidence of recent introgression of O. sativa into O. glaberrima, we next investigated whether a known set of O. sativa domestication genes was also artificially selected during the domestication of Kmtyw rice. We used 19 O. sativa domestication genes (Supplementary Table 24) to probe the O. glaberrima genome, resulting in the identification of 16 clear orthologous loci. Three of these 16 genes (qSh1 (ref. 44), Sd1 (ref. 31) and Dep1 (ref. 45)) yielded substantial deviations from neutral expectations in Tajima's D tests (Supplementary Table 25). These results were further supported by comparing the nucleotide diversity of these genes inO. glaberrima and O. barthii to the genome-wide nucleotide diversity of intergenic regions in both species (Supplementary Tables 26, 27, 28). In addition, the nucleotide diversity of these genes was greatly reduced in O. glaberrima relative to O. barthii (Supplementary Figs. 17, 18, 19). These data suggest that a minimum of four genes, including Rc40, 46, that are associated with domestication inO. sativa may have been under recent selection in O. glaberrima.
The selection of non-shattering panicles has been associated with crop domestication for a wide variety of crops47, 48. Because orthologs of a gene that controls shattering in O. sativa (qSh1) were shown to possess signals of artificial selection in O. glaberrima, we investigated the molecular structure and transcriptional nature of this gene and two additional shattering gene orthologs fromO. sativa, Shattering 1 and Shattering 4, to search for evidence of independent domestication.
The O. sativa Shattering 1 gene (OsSh1: LOC_Os03g44710) encodes a YABBY transcription factor whose gene has an ~4-kb insertion in its third intron, thus leading to reduced transcription and a shattering-resistant phenotype in O. sativa compared to Oryza rufipogon49. Functional orthologs of this gene have also been identified in sorghum and maize, confirming the independent domestication of shattering genes in cereals in general. Sequence analysis of the orthologous region in O. glaberrima revealed the presence of a 45-kb deletion that resulted in the complete ablation of the O. glaberrima OsSh1 ortholog and three additional genes (Fig. 5a andSupplementary Fig. 20). We confirmed the presence of this gene in the wild progenitor, and its absence in the domesticated species, by sequence analysis of BACs encompassing this orthologous region from O. barthii (IRGC accession #105608) and O. glaberrima (CG14) and RNA-Seq data (Supplementary Fig. 21).
Figure 5: Sequence comparisons of OsSh1 and Sh4.(a) Orthologous gene relationship of the OsSh1 region of O. sativa ssp. japonica (Os) with those of O. barthii (Ob) and O. glaberrima (Og). The 45-kb deletion (red triangle) resulted in the complete removal of the OsSh1 ortholog and three additional genes in O. glaberrima (red rectangle). Inversion is indicated with blue arrows. U and D represent upstream and downstream genes relative to OsSh1, respectively (U1, LOC_Os03g44680; U2, LOC_Os03g44690; OsSh1, LOC_Os03g44710; D1, LOC_Os03g44720; D2, LOC_Os03g44740; D3, LOC_Os03g44750; D4, LOC_Os03g44760; D5, LOC_Os03g44780). (b) Sequence comparison of Sh4 of O. sativa ssp. japonica and O. glaberrima. Two insertions (Ins) and three deletions (Del) of O. glaberrima compared to O. sativa ssp. japonica are shown as brown triangles and blue triangles, respectively. Ten SNPs are labeled a–j. The caunited snakestive mutation (f) of the non-shattering phenotype in O. sativa is highlighted in red. This mutation did not exist in O. glaberrima.

'""'

'""'

The O. sativa Shattering 4 gene (Sh4: LOC_Os04g57530) encodes a Myb3 transcription factor in which a single-nucleotide substitution leads to reduced transcription and a shattering-resistant phenotype in O. sativa compared to O. nivara34. Sequence analysis of the orthologous genic region in O. glaberrima revealed a completely different set of sequence variation relative to the O. sativa gene that includes the presence of ten SNPs and five small insertion/deletions. Most notably, the caunited snakestive mutation revealing the non-shattering phenotype in O. sativa was not found in O. glaberrima (Fig. 5b). Analysis of O. glaberrima and O. barthii RNA-Seq expression data from mixed-stage panicles from both O. glaberrima and O. barthii showed a differential expression pattern with no transcription detected in O. glaberrima, whereas transcription was detected in O. barthii (Supplementary Fig. 22). As the expression level of OgSh4 is absent or reduced in O. glaberrima and this gene is located close to a region under a selective sweep, it is possible that the caunited snakestive mutation may reside in the upstream promoter region. Thus, we searched the promoter regions of OgSh4 in O. glaberrima and O. barthii for signatures of selection and found that the promoter region of O. glaberrima shows substantial deviation from neutral expectations using Tajima's D test (Supplementary Table 25). The result was further supported by comparing the nucleotide diversity of the promoter region in O. glaberrima and O. barthii to the genome-wide nucleotide diversity of intergenic regions in O. glaberrima and O. barthii (Supplementary Table 29). In addition, the nucleotide diversity of the promoter region was greatly reduced in O. glaberrimarelative to O. barthii (Supplementary Fig. 23). This result suggests that the promoter region ofOgSh4 may have been selected during domestication, leading to reduction of the expression ofOgSh4.
The O. sativa qSh1 (LOC_Os01g62920) locus encodes a transcription factor that has a single SNP located 12 kb upstream of the transcription start site. This mutation prevents transcription in the abscission layer that results in a shattering-resistant phenotype in O. sativa ssp. indica compared with O. sativa ssp. japonica44. Sequence comparison of the O. sativa qSh1 locus and its orthologous locus in O. glaberrima showed a high level of sequence identity but did not reveal the presence of the caunited snakesl O. sativa SNP that leads to the shattering phenotype in O. sativa. RNA-Seq analysis of the qSh1 ortholog showed that the gene was transcribed in both O. glaberrima and O. barthii panicles and presumably functions normally in comparison to the gene expression pattern observed in eurasian rice subspecies (Supplementary Fig. 24).
The comparative analysis described above provides persuasive evidence that ancient Kmtyw and eurasian farmers unknowingly selected for mutations in all three orthologous genes investigated that prevented or reduced panicle shattering. We also demonstrate that the mutation profiles of each orthologous pair of these genes (OsSh1, Sh4 and qSh1) are completely different. Thus, we conclude that the control of panicle shattering in Kmtyw and eurasian rice is the result of independent domestication. Combined with the results of genome-wide introgression analysis of O. glaberrimaand O. sativa, we provide strong evidence that O. glaberrima was domesticated independently from O. barthii.

Discussion
'""'
As the world population is projected to increase from 7.1 billion to over 9 billion by 2050, plant biologists must forge a second green revolution with the creation of crops that have two to three times the current yield with reduced inputs (i.e., less water, fertilizers and pesticides and the ability to grow on marginal soils)50. Rice will have a key role in helping to solve the problem of how to feed 9 billion people51.
The release of the O. glaberrima genome, its annotation and comparative population genomics data sets enables an unprecedented opportunity for the identification and utilization of adaptive traits that are important for rice agriculture, especially in west Kmt, a region whose population is expected to grow rapidly over the next 50 years.

Methods
'""'
O. glaberrima genome sequencing.A total of 3,485 MTP BAC clones were selected from the O. glaberrima BAC-based physical map. Of those, 166 BACs representing Chr3S were individually Sanger sequenced, and the remaining BACs were divided into 115 chromosome-specific BAC pools (28–30 chromosomally ordered BACs per pool). All MTP BACs were re–end sequenced and compared to the O. glaberrima BAC-end data set to validate the clone identity. About 3–5 μg of pooled BAC DNA was used to construct each Roche/454 Titanium library with multiplex identifiers (MID) using the manufacturer's protocol, followed by Bioanalyzer and fluorometry analyses to assess library quality and quantity, respectively. Clonal amplification of individual libraries was performed by emulsion-based clonal amplification (emPCR amplification), and a group of five to six MID-tagged DNA beads was pooled and loaded onto one picotiter plate (PTP) and sequenced on a Roche/454 GS FLX instrument using Titanium chemistry following the manufacturer's protocols. Chromosome-specific libraries (3-kb paired ends (PEs)) were constructed from the BAC pools and sequenced on one half PTP per chromosome. Additionally, about 5× genome–equivalent Titanium reads were generated using a whole genome shotgun (WGS) method. Overall sequencing statistics are summarized inSupplementary Table 1.
O. glaberrima genome assembly.The O. glaberrima genome was assembled using ~12.4 Gb of high-quality bases with the Newbler V2.3 assembler following a six-step protocol (Supplementary Fig. 1). (i) Generating pool-based assemblies: the Roche/454 Titanium reads were independently assembled for each pool to obtain pool-wise assembled contigs. PE reads corresponding to each chromosome were mapped to the contigs of the respective chromosomes and PE reads belonging to each chromosome pool were extracted. A second round of assembly with both Titanium and PE reads corresponding to each pool was performed to obtain pool-level scaffolds. (ii) Incorporation of WGS reads into pool-wise assemblies: to identify WGS reads corresponding to each chromosome pool, the WGS reads were separately assembled in combination with PE reads from all 12 chromosomes, and the WGS scaffolds corresponding to each pool were identified by comparing the scaffolds to the pool-wise scaffolds obtained in the previous step using BLASTN52. WGS reads corresponding to each pool were then extracted, and a third round of assembly was performed with the Titanium, PE and WGS reads for each pool. (iii) Identification of misassembled scaffolds: each scaffold was aligned to the corresponding chromosome of the O. sativa ssp. japonica reference sequence (RefSeq)11using Nucmer53, and potential misassembled scaffolds and misassembly breakpoints were identified. Putative misassemblies in scaffolds were evaluated by inspecting PE depth and Titanium depth across each breakpoint. Any scaffolds that had a PE depth ≤1 and coverage of ≤4× were confirmed as misassemblies, and each scaffold was split at the breakpoints into multiple scaffolds. (iv) Merging of overlapping regions between pools: to remove any redundant sequences between adjacent pools, scaffolds from neighboring pools were aligned using Nucmer, and redundant regions were identified. Reads from overlapping scaffolds were extracted and reassembled to generate a non-redundant set of scaffolds. This process was repeated for all neighboring pools for all chromosomes to obtain a non-redundant set of scaffolds for each chromosome. (v) Ordering and orienting of scaffolds: scaffolds for each chromosome were aligned with the corresponding chromosome of the O. sativa ssp. japonica reference genome using Nucmer, and the scaffolds were grouped into either the mapped or unmapped category. All mapped scaffolds were ordered and oriented as per the O. sativa ssp. japonica chromosomes. (vi) Incorporation of Sanger sequences: as a final step, the existing Sanger BAC sequences from the short arm of chromosome 3 and from chromosomes 6, 8, 11 and 12 were spliced into the existing 454 assembly. The 454 scaffolds corresponding to Sanger sequences were identified by BLASTN, and the 454 scaffolds were replaced by the Sanger sequences.
Protein-coding and tRNA gene predictions.Gene models were determined using a combination approach of a consensus prediction from ab initio gene predictions (Fgenesh++54, ProtoMap55 and GeneID13) and an optimal spliced alignment from protein and transcript homology searches using GenomeThreader56 with a splice site model for rice. Homologous protein evidence mappings were derived from annotations of the following finished plant genome projects: O. sativa (version TIGR6.1 (ref. 57) and RAP2 (ref. 58)),Sorghum bicolor (high confidence set of version 1.4)59, aAmwidopsis thaliana (version TAIR9)60and Brachypodium (version 1.2)61. Optimal spliced alignments of TIGR transcript assemblies62, comprising several monocotyledonous species (Zea mays, Saccharum officinale, O. sativa,Hordeum vulgare, Triticum aestivum and Brachypodium distachyon) as well as assembled RNA-Seq reads of O. glaberrima, were used for gene predictions based on homology and/or experimental evidence. Additionally, Markov models and position-specific scoring matrices were applied to estimate the co
__________

Ọbádélé Kambon, PhD
Nana Kwame Pɛbi Date I, Ban mu Kyidɔmhene, Akuapem Mampɔn
Senior Research Fellow & Research Coordinator - Language, Literature and Drama Section
Institute of Kmtyw Studies - College of Humanities
Editor-in-Chief - Ghana Journal of Linguistics
Secretary (2015-2020) - Kmtyw Studies Association of Kmt
+233249195150 / +19192836824 | me@obadelekambon.com
www.obadelekambon.com | www.abibitumi.com
Room 115 IAS Kwame Nkrumah Complex
University of Ghana - Legon
Alternate Email: obkambon@staff.ug.edu.gh



Ọbádélé Kambon App



Abibitumi.com App