Sequins™ WGS Bioinformatics Resources

Sequins™ WGS Bioinformatics Resources
Sequins™ WGS Bioinformatics Resources

Sequins are synthetic DNA standards designed to mirror natural genomic features, enabling enhanced quality insights and improved calling confidence of whole genome sequencing (WGS) data.

As an NGS spike-in, Sequins provide internal controls for benchmarking performance across library preparation, sequencing, and data analysis. When using Sequins in WGS, key bioinformatics considerations include accurate alignment of synthetic sequences, variant calling within the spike-in reference, and comparison of observed versus expected metrics to assess sensitivity, specificity, and dynamic range.

This section provides guidance to support effective use of Sequins in your WGS workflows. If you require any further assistance, please email support@sequins.bio.

 

WGS Bioinformatics Workflow Overview

The diagram illustrates a five-step workflow for integrating Sequins into WGS bioinformatics pipelines, enabling robust performance benchmarking and variant calling accuracy assessment.

This structured process supports both Docker-based and independent toolchain implementations and is tailored to be flexible for diverse research and clinical applications using Sequins in WGS.

 

Step 1: Create Sequins Augmented Reference file

Combine the standard reference genome (e.g., GRCh38) with the Sequins decoy chromosome to generate an augmented reference. This step is performed once and used throughout the workflow.

 

Step 2: Read Alignment

Sequence reads (FASTQ) are aligned to the augmented reference genome using a tool such as BWA, producing a BAM file that contains both sample and Sequins reads.

 

Step 3: Calibrate Sequins Reads

Sequins are typically sequenced at higher coverage than the sample. To ensure fair comparison, the Sequins regions in the BAM are downsampled to match the coverage of their corresponding human genome locations. This calibrated BAM file better reflects variant detection sensitivity in the actual sample.

 

Step 4: Variant Calling

Variants are called separately on the calibrated BAM for Sequins and the sample. This step produces two VCF files: one for the Sequins controls and one for the sample.

 

Step 5: Performance Evaluation (Control Analysis)

By comparing the known variant truth set of Sequins (VCF Sequins) to the variant calls obtained from the pipeline, users can generate precision, sensitivity, and F-measure metrics using RTG Tools. This enables identification of false positives and negatives, and provides a clear benchmark of the pipeline’s performance.

 

WGS Sequins Tutorial

Here you can find an example workflow for processing data that have had Sequins spiked in. It highlights the places where Sequins specific steps must be taken, and it is not intended as an example of a production workflow.

Learn more about WGS Sequins tutorial

Learn more about WGS Sequins tutorial

Frequently asked questions

How does using Sequins impact my WGS bioinformatics workflow?

There is minimal impact to your bioinformatics workflow. There are two (2) minor additions to a standard WGS bioinformatics workflow:

  1. An augmented reference genome file (combination of standard human genome reference file and Sequins reference file).
  2. Integration of Sequins’ read downsampling software (sequintools) into your bioinformatics workflow.

How do I create the augmented reference genome file for my WGS bioinformatics workflow?

Simply combine your desired human reference genome file with the Sequins reference file.  One method is to use the cat command on Unix.

Where do I get the Sequins reference file?

The Sequins team will provide you with the Sequins decoy reference file.

Why align to an augmented reference?

So Sequins reads align to the Sequins reference portions of the augmented reference, while sample reads align with the human reference genome portion of the augmented reference.

How do I integrate sequintools into my WGS bioinformatics workflow?

The Sequins molecules added to your sample are short sequences of nucleic acids, which will naturally accrue higher cover than other relative-sized regions in the rest of the sample. As a result, it is important to downsample the Sequins reads to be more relative to the number of reads in the rest of the sample. The downsampling step is done before variant calling. To accomplish this, simply add the sequintools docker into your bioinformatics pipeline. The sequintools docker should be compatible with standard scientific workflow managers, such as Nextflow.

Where do I get sequintools (Sequins software)?

sequintools is open source software, and it is available on GitHub: (https://github.com/sequinsbio/sequintools).

What aligners are compatible with Sequins?

Sequins will minimally impact your bioinformatics workflow and are compatible with any aligner.

What variant callers are compatible with Sequins?

Sequins will minimally impact your bioinformatics workflow and are compatible with any variant caller.

Why do the Sequins align next to each other, are they designed in one contiguous molecule?

Sequins are individual molecules, not one continuous molecule. Placing all sequences on the decoy FASTA file that is provided is purely a convenience and the Sequins can be analyzed by aligning to the individual sequences if you prefer.

Why do I see gaps between my Sequins regions?

Because the decoy FASTA is a synthetic construct and the individual Sequins are not really one continuous molecule.

Why do I see drop off in coverage at the terminal regions of each Sequins?

This is just a natural effect of sequencing to the end of a molecule. As you reach the end of a molecule there are fewer positions a read can start or end so there is a natural drop off in coverage.

I’m using Dragen, how do I incorporate Sequins?

The preferred approach is to add the Sequins decoy chromosome to the human reference FASTA and generate a DRAGEN reference hash from the combined reference. This allows human and Sequins reads to be aligned in the same analysis while keeping the standard downstream interpretation of the human genome unchanged.

If modifying the DRAGEN reference is not possible, the reads reported as unmapped from the standard human reference genome analysis can be extracted and aligned separately to the Sequins reference.

However, the unmapped-read approach should be validated as depending on how DRAGEN classifies and exports reads, not all Sequins-associated reads may be retained in the unmapped output. If Sequins coverage is unexpectedly low or incomplete, the relevant read IDs should be traced back to the original DRAGEN BAM and their alignment status, mapping quality and SAM flags reviewed.

Why can sample-matched calibration fail for mitochondrial Sequins regions in the WGS Core Control Set?

Native mitochondrial DNA coverage can vary substantially between samples because of differences in mitochondrial content and mtDNA copy number. In some WGS datasets, native chrM coverage may therefore exceed the coverage of the corresponding mitochondrial Sequins regions.

sequintools calibrate can downsample Sequins reads but cannot increase their coverage. If native chrM coverage exceeds the coverage of the corresponding mitochondrial Sequins regions (SG_000000038–SG_000000041), sample-matched calibration will stop with an error.

This does not necessarily indicate poor recovery or mapping. It means that these mitochondrial regions cannot be calibrated to the native chrM depth in that dataset.

In such cases, the following options are available:

  1. Exclude the mitochondrial Sequins from sample-matched calibration.
    Remove SG_000000038–SG_000000041 from the decoy chromosome BED file and calibrate the remaining Sequins regions normally. If mitochondrial Sequins reads are required in the final BAM or CRAM file, they can be extracted before calibration and merged back afterwards.

The mitochondrial Sequins regions may still be assessed as QC markers for recovery, mapping and coverage consistency, but should not be interpreted as sample-matched controls for native chrM depth.

  1. Use flat-target calibration.
    Where appropriate for the analysis, setting a fixed target coverage, such as -f 40, avoids this specific issue because calibration does not attempt to match native chrM coverage.

This interpretation is specific to mitochondrial regions. Under the recommended spike-in and workflow conditions, nuclear Sequins coverage is expected to exceed the coverage of the corresponding native regions. The same error for a nuclear region should therefore be investigated as a possible input, workflow or configuration issue.

Request access to tutorial datasets

"*" indicates required fields

Area(s) of Interest*
You agree to our friendly Privacy Policy
Table of Contents

Unlock valuable insights by gaining access to our comprehensive dataset. Click below to get started.