Genetic Variation & Association Analysis

Variant analysis built for African genomes — where reference bias, population structure and under-representation in annotation databases all matter.

African populations carry more genetic diversity than the rest of the world combined, and are the most under-represented in the reference panels and annotation databases that variant pipelines depend on. A pipeline tuned on European cohorts will systematically miss and misclassify variants in African samples.

We build that into how we work: appropriate reference panels, population-specific allele frequency filters, and honest reporting of where annotation coverage is thin.

Data We Accept

  • Raw FASTQ files from whole genome, exome or targeted panel sequencing
  • Aligned BAM or CRAM files
  • VCF or gVCF files for annotation, filtering or joint analysis
  • Genotyping array data (PLINK binary, Illumina final report)
  • Pedigree files and phenotype tables

Questions We Answer

  • Which variant is causing this family's disease?
  • Which loci are associated with my trait or outcome?
  • How is my cohort structured by ancestry, and does it confound my analysis?
  • Are the variants I found real, or artefacts of coverage and reference bias?
  • How does my population compare to reference panels?

What We Do

Each project uses the subset of these that your research question requires.

01

Alignment and Variant Calling

GATK best practices or DeepVariant for germline SNVs and indels, joint genotyping across cohorts, and structural variant and copy number calling where the data supports it.

02

Annotation and Prioritisation

Functional annotation with VEP, ANNOVAR or SnpEff against ClinVar, gnomAD, dbNSFP and African-specific frequency resources, with variants prioritised by consequence, frequency and inheritance model.

03

Rare Disease and Trio Analysis

De novo, recessive and compound heterozygous filtering across trios and families, with phenotype-driven ranking using HPO terms and Exomiser.

04

Genome-Wide Association Studies

Quality control, population stratification via principal components, imputation against appropriate reference panels, and association testing with PLINK, REGENIE or SAIGE — including mixed models for related samples.

05

Population Genetics

Admixture and ancestry inference, runs of homozygosity, fixation index and selection scans for population and evolutionary studies.

What You Receive

  • Annotated and filtered VCF plus a readable variant table
  • Prioritised candidate variant list with evidence for each call
  • Manhattan, QQ and regional association plots for GWAS
  • Population structure plots and ancestry estimates
  • Coverage and QC report per sample
  • Full pipeline code and reference versions used

Tools We Use

  • BWA-MEM2, GATK4, DeepVariant
  • bcftools, samtools, Hail
  • VEP, ANNOVAR, SnpEff, Exomiser
  • PLINK 2, REGENIE, SAIGE
  • ADMIXTURE, EIGENSOFT, selscan
  • Manta, CNVnator, DELLY

Typical turnaround: 3–5 weeks for exome or genome cohorts; 2–4 weeks for GWAS from array data

Indicative price: From $99 per sample for exome calling and annotation at cohort scale

Reduced rates are available for students and researchers at African public institutions. Every project is quoted in writing before work begins.

Services are provided for research purposes only. They are not intended for clinical diagnosis, treatment decisions or individual health assessment. See how it works, data submission guidelines and what you receive.

Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827