Biodiversity informatics · DNA barcoding infrastructure · machine learning for taxonomic inference · knowledge systems for life science
I design and build the data systems that let thousands of researchers turn DNA sequences into knowledge about life on Earth. I am the architect of BOLD, the Barcode of Life Data System; the BIN framework for DNA-based species registration; and mBRAVE, a platform for high-throughput metabarcoding analysis. My current work is on AI-queryable knowledge graphs and governed AI for biodiversity science.
My work sits at the intersection of biodiversity science, large-scale data engineering and machine learning. The recurring question is how to make biological data at planetary scale usable: interoperable across institutions, computable by models, and trustworthy enough for regulatory and forensic use.
Architecture of federated, multi-tenant platforms for specimen, sequence and observation data; provenance and governance models; data standards such as the Barcode Core Data Model (BCDM) and its mappings to Darwin Core and MIxS.
Algorithms that assign sequences to species without complete reference libraries: the BIN clustering framework, GPU-accelerated probabilistic classification (PROTAX-GPU), and multimodal insect datasets for representation learning (BIOSCAN-1M / 5M).
High-throughput pipelines for environmental DNA and bulk-sample sequencing (mBRAVE), and statistical methods that extend inference from common to rare species, such as CORAL, applied to a quarter-million Malagasy arthropod taxa.
Domain-specific knowledge graphs that make biodiversity data queryable by language models with provenance intact; on-premises, auditable AI pipelines for government and regulatory programmes.
Each platform below is in production and used by the global DNA barcoding community. Together they form a stack from raw sequence to institutional knowledge.
Citation counts from Google Scholar. Full list on Google Scholar.
Public repositories from github.com/DNAdiversity. This list refreshes from the GitHub API when the page loads; the entries below are the fallback.
Awarded by the Global Biodiversity Information Facility (GBIF) for the development of BOLD, "a major and innovative landmark in bringing genomic data on biodiversity to research." First Canadian recipient.
BOLD is listed by the U.S. National Institute of Standards and Technology as a species-identification resource in its forensic database guidance.
Member, convened by Canada's Office of the Chief Science Advisor.
Co-lead of Working Group 3.1, responsible for the informatics backbone of the International Barcode of Life project.
I am open to academic collaborations, advisory roles and conversations about biodiversity data infrastructure, taxonomic ML and governed AI.
sujeevan@ratnasingham.com
Guelph, Ontario, Canada