Cacao genebanks preserve the diversity needed to breed trees for changing climates, emerging diseases, and more resilient production. However, historical mislabeling, synonymy—one genotype recorded under several names—and homonymy—different genotypes sharing one name—can undermine breeding and field recommendations. High-density genotyping can resolve these problems, but routine use may demand more money, laboratory capacity, and data handling than many programs can support. Existing low-density panels are cheaper, yet their markers are often selected without a unified method for comparing identification power, population structure, and trait relevance. Because of these challenges, deeper research is needed to develop compact genotyping systems that preserve the information most useful for germplasm management and breeding.
Researchers from the Agricultural Research Service of the United States Department of Agriculture (USDA-ARS) published (DOI: 10.1093/hr/uhag106) the study on March 19, 2026, in Horticulture Research. The team represented the Sustainable Perennial Crops Laboratory and the Environmental Microbial and Food Safety Laboratory in Beltsville, Maryland, and the Grape Genetics Research Unit in Geneva, New York. Using cacao collections from Trinidad and Puerto Rico, they developed and evaluated an information-theoretic method that reduced a dataset containing 536 single nucleotide polymorphisms (SNPs) to a 32-marker barcode while balancing identification, ancestry screening, and agronomic information.
The researchers built CacaoCipher using 346 accessions from the International Cocoa Genebank, Trinidad (ICGT), and tested cross-environment consistency with a field trial at the USDA-ARS Tropical Agriculture Research Station (TARS) in Puerto Rico. A stepwise selection procedure favored informative SNPs while controlling missing data and linkage disequilibrium (LD), or non-random association among variants. Reducing 536 SNPs to 32 cut the marker count by 94% while preserving coarse genetic relationships; a Mantel test comparing the reduced and full genetic-distance matrices yielded a correlation of 0.84. Simulations indicated that the barcode retained more than 95% correct identification even when the per-locus error rate reached 5%. The pure-identification panel also outperformed randomly chosen 32-marker sets in ancestry assignment, although it was suitable only for coarse-to-intermediate screening. A trait-enriched version increased marker–trait correlations by 39% for pod index—the number of pods needed to produce 1 kilogram of dried beans—and 65% for mature fruit-surface anthocyanin intensity, with only a modest loss of genetic distance. Among 24 accessions shared between the Trinidad and Puerto Rico datasets, barcode coordinates were associated with yield, total pod number, pod index, and dry bean weight. The analysis also identified a near-identical group centered on NA326 containing 148 accessions, suggesting extensive duplication or mislabeling.
The authors said the study changes the question from how many markers can be measured to how a limited marker budget should be spent. They said CacaoCipher demonstrates that a compact barcode can support identity checks while still carrying useful ancestry and trait-aligned signals. In their view, identity and agronomic screening are not competing endpoints but adjustable design priorities. The authors also emphasized that the barcode is intended to complement high-density genotyping, giving curators a rapid first-line tool while reserving denser panels for pedigree reconstruction, fine-scale selection, and other tasks requiring greater resolution.
CacaoCipher could help genebanks authenticate accessions, detect duplicates, screen broad ancestry, and prioritize material for more detailed testing. The estimated 37-bit capacity was allocated heuristically, with about 58% supporting unique identification, 32% reflecting ancestry structure, and 10% aligned with yield and pod efficiency. These proportions are design guides rather than independent biological compartments, because the information targets overlap. The panel is not yet validated for degraded samples, mixed-origin products, adulteration detection, or definitive parentage analysis. Future work should test it across genotyping platforms, larger independent populations, and controlled traceability conditions. More broadly, the genomic bit budget provides a transferable way to design low-cost, information-efficient barcodes for cacao and other clonally propagated crops.
###
References
DOI
10.1093/hr/uhag106
Original Source URL
https://doi.org/10.1093/hr/uhag106
Funding information
This work was supported by United States Department of Agriculture, Agricultural Research Service, In-House Projects 8042-21220-258-000-D and 8072-41000-113-000-D.
About Horticulture Research
Horticulture Research is an open access journal of Nanjing Agricultural University and ranked number one in the Horticulture category of the Journal Citation Reports ™ from Clarivate, 2023. The journal is committed to publishing original research articles, reviews, perspectives, comments, correspondence articles and letters to the editor related to all major horticultural plants and disciplines, including biotechnology, breeding, cellular and molecular biology, evolution, genetics, inter-species interactions, physiology, and the origination and domestication of crops.