top of page

Our Science

We sequence DNA from the product itself

When a sample arrives at our laboratory, we extract DNA directly from the material, prepare it for sequencing, and run it on a nanopore sequencing platform that reads individual DNA molecules in real time. The sequencer generates long reads of genetic data, which are then classified against curated reference databases to determine what species are present in the sample, their relative abundance in the sequencing data, and whether anything unexpected is there

.
The workflow, from sample receipt through extraction, library preparation, sequencing, bioinformatic classification, and report generation, is designed to be performed entirely in-house in our BSL-2 (biosafety level 2) laboratory in Tampa, Florida.

Sequencing platform

We sequence on a long-read nanopore platform from Oxford Nanopore Technologies (ONT). Where short-read sequencers produce fragments of 150 to 300 base pairs, nanopore sequencing reads individual DNA molecules as they pass through a protein pore, generating reads that routinely span tens of thousands of bases and can exceed 100,000. The sequencer streams data in real time, which means results can be monitored as they accumulate and analytical decisions can be made during the run.


Long reads are valuable for product authentication because they provide more taxonomic information per read, span repetitive genomic regions that short reads struggle to resolve, and produce classification results with higher confidence at lower read depths. For projects that require short-read sequencing, such as high-accuracy variant calling or applications where established workflows mandate it, we maintain relationships with core sequencing facilities and coordinate those runs as part of a complete project.

Sample preparation

DNA extraction is the step that determines what the sequencer has to work with, and different product types require different approaches. Animal tissues and dried botanical supplements generally yield high-quality DNA with standard column-based extraction kits. Processed liquids, such as honey, mead, wine, and kombucha, present a greater extraction challenge because the DNA is dilute, degraded, or bound up in a complex matrix that requires concentration before extraction. Our current approach for these matrices uses vacuum filtration through membranes to capture cells, pollen, or cell-free DNA before lysis, a method validated in published environmental DNA research and adapted in our laboratory for food and beverage applications.


After extraction, every sample is quantified using fluorometric assay (Qubit) to measure the actual concentration of usable DNA. When the extraction matrix necessitates it, spectrophotometric purity ratios confirm whether the extract is clean enough for library preparation, and, for samples where fragment length is critical to the application, electrophoretic fragment analysis provides a detailed size distribution to confirm that the DNA is suitable for the intended library preparation and sequencing approach.


Library preparation converts extracted DNA into a form the sequencer can read. We currently use two primary approaches, depending on the project. Rapid library preparation uses a transposase enzyme to simultaneously fragment DNA and attach sequencing adapters, which is well suited for species identification work where speed matters. Ligation-based preparation preserves longer fragments and produces higher-quality data for applications like whole genome sequencing or structural analysis at the cost of a longer preparation time. Both approaches support multiplexing many samples on a single flow cell, which allows for increased throughput, lower processing time, and more accessible pricing for our clients. 

Bioinformatic classification

Raw sequencing data is basecalled (converted from electrical signal to nucleotide sequence) using ONT's neural network basecaller. We select the highest accuracy basecalling model appropriate for each run to maximize per-read quality and efficiency, and we upgrade to newer models as they're available.


Basecalled reads are classified using a combination of tools chosen to match the analytical question. K-mer based taxonomic classifiers compare short subsequences from each read against indexed reference databases and assign taxonomic labels based on the best match. Alignment-based tools map full reads against curated reference genomes for finer resolution when k-mer classification reaches genus level but species-level discrimination requires more precision. For samples that require additional confirmation, ambiguous classifications are flagged for manual review and verification against the full NCBI nucleotide database. The goal is to report species-level identifications where the data supports them and to be explicit about the taxonomic level at which confidence is justified when it does not.


Our bespoke reference database architecture is built in tiers designed for the specific demands of product authentication. The organellar backbone contains over 33,000 mitochondrial and plastid genomes from NCBI RefSeq, covering a broad base of animal, plant, and fungal species. On top of that foundation, we are building curated nuclear genome assemblies for priority species across our core product categories: commercial seafood, livestock, medicinal and culinary botanicals, edible fungi, common adulterants, species of conservation concern, and other species as needed for clients we work with. Additional reference libraries for bacteria, fungi, plants, and human sequences provide context for complex metagenomic samples where background organisms and contaminants need to be identified alongside the target material. This architecture is designed to grow and be restructured as our validation work matures and as the demands of new product categories shape what the database needs to contain.

 

For clients who need the ability to verify a proprietary cultivar, a specific strain, or a commercial variety that public reference databases do not cover, we provide custom reference database development. We sequence the client's authenticated material, build a validated reference entry, and integrate it into our classification system so that the client's specific verification question can be addressed within our classification system. This capability is particularly relevant for clients developing proprietary varieties, specialty agricultural products, or branded ingredients where public reference databases have no coverage.

Whole genome and metagenomic sequencing

Most commercial DNA testing for product authenticity uses targeted PCR or other amplification based approaches. This means amplifying specific genetic markers (such as COI for animals, ITS for fungi, or rbcL and matK for plants) and comparing those amplicons against reference databases. This is effective when the question is narrow and the expected answer falls within a known set of species.


Our laboratory is built on whole genome and metagenomic sequencing as a foundation, which means we sequence all of the DNA recovered from a sample rather than targeting predetermined markers. This approach is designed to answer a broader range of questions. It can identify species that a targeted panel was never designed to detect, including unexpected adulterants and contaminants. It provides data on the full biological composition of a sample, rather than a binary presence or absence call for a predefined list. And, it generates a richer dataset that can be re-analyzed as reference databases improve, without requiring the client to submit a new sample.


For sample types where background DNA dominates the material of interest (such as fermented beverages where yeast and bacterial DNA vastly outnumber the plant-derived DNA that answers the authentication question), we are developing targeted enrichment workflows that increase the proportion of informative reads in a sequencing run. These approaches use real-time or pre-sequencing strategies to focus sequencing capacity on the organisms and genomic regions relevant to the authentication question, improving sensitivity for low-abundance targets.


Targeted approaches, including PCR based methods, remain part of our toolkit for cases where a focused answer is what the question calls for, or where sample condition favors fragment amplification over long read sequencing.

What we report

Every certificate of analysis documents the sequencing platform, library preparation method, basecalling model and version, reference database version, and bioinformatic pipeline used to produce the result. Species identifications are accompanied by supporting data including classification scores and read depth, and the report distinguishes between high-confidence identifications, lower-confidence detections that may warrant follow-up, and results where recoverable DNA was insufficient for classification.


We report to you what we find, including species that are expected, species that are unexpected, and cases where recovery falls below what a client anticipated. Reports are designed to be readable by non-scientists while containing enough methodological detail that a qualified reviewer can evaluate how we arrived at our conclusions and assess the methodology on its merits. Clients who require the underlying sequence data for independent re-analysis can request it.

What DNA sequencing can determine, and where it has limits

Species-level identification is robust for organisms that are well represented in reference databases and from which sufficient DNA can be recovered. Most animal tissues, dried botanicals, whole spices, raw and minimally processed foods, and natural textiles are well suited to this kind of analysis.


Detection of undeclared components is effective for biological materials present above our reporting threshold. If a supplement contains a species absent from the label, or if a seafood product contains a species different from the one claimed, sequencing is designed to identify it as long as the DNA is recoverable and the species exists in reference databases.


Subspecies, cultivar, and variety-level identification is possible for species where reference data at that resolution is available. Some distinctions require specialized marker panels or deeper sequencing than a standard authentication run provides, and we scope those projects accordingly.


Geographic origin determination through DNA alone is possible for some species and populations, particularly where genetic structure varies across geography, but is unreliable for commodity crops with genetically uniform global populations. We are straightforward about where this capability applies and where it does not.
Heavily processed products (refined oils, highly filtered liquids, extensively heat-treated materials) may yield degraded or undetectable DNA. When a sample type is likely to present extraction challenges, we communicate that upfront and adjust expectations and pricing accordingly.

Ongoing development

Our reference databases, classification pipelines, and sample preparation protocols are under continuous development. We add new species and new reference assemblies as projects demand them and as new genome data becomes available. We refine extraction protocols for challenging matrices as we work with more sample types. And we are building the infrastructure to make our verification results more durable and independently confirmable over time, work that is described on our Philosophy page.


The science underneath product authentication is well established. The work we are doing is applying it systematically, building the databases and protocols and reporting standards that turn a research capability into a reliable service, and making that service available to the people who need it most. If you have questions about how our analytical workflow applies to your product or sample type, we want to hear from you.

bottom of page