Case Studies

Worked examples showing how a DataCore project runs, from first enquiry to handover.

These are illustrative examples, not client case studies. DataCore Analytics is a young company and we will not present invented client work as real. Every example below describes the kind of project we take on, the decisions involved and what the client receives — using realistic parameters drawn from the study designs our consultants work with.

As real engagements complete and clients agree to be named, this page will be replaced with their work.

Pathogen Genomics

Genomic Surveillance for a National Reference Laboratory

A national public health laboratory sequencing respiratory pathogen samples weekly needs lineage assignment and outbreak detection with turnaround fast enough to inform response.

What arrived

Roughly 200 amplicon-sequenced genomes per week arriving from Illumina and Nanopore runs, with no in-house bioinformatician. Consensus genomes were being generated inconsistently and submissions to public repositories had stalled.

What we did
  • Established a standing weekly batch agreement rather than per-project quoting
  • Built a single reproducible consensus pipeline covering both Illumina and Nanopore inputs, with primer scheme handling and per-genome quality flags
  • Automated lineage assignment and produced a submission-ready package for GISAID and ENA
  • Set up a Nextstrain build the laboratory hosts and updates itself
  • Delivered three training days so laboratory staff could run the pipeline and interpret the output
What the client received
  • Weekly batch report circulated to the health ministry within 48 hours of run completion
  • Consensus FASTA with quality flags and a repository submission package
  • Transmission cluster alerts where sequences fell within a defined genetic distance
  • A pipeline the laboratory now runs without us, with us available for escalation

Timeline: Standing partnership; 48-hour turnaround per batch

Clinical & Epidemiological Data

Underpowered Cohort: Telling the Client Before the Reviewer Does

A research group asked for differential expression analysis on a cohort collected over three years, expecting a manuscript. The analysis showed the study could not answer the question as designed.

What arrived

RNA-Seq on 18 samples across three clinical groups, collected in two waves separated by 14 months, with extraction batch perfectly confounded with clinical group.

What we did
  • Ran quality control first and modelled batch structure before touching differential expression
  • Established that batch and clinical group could not be separated statistically — any apparent difference was uninterpretable
  • Reported this within the first week rather than producing a gene list that would not survive review
  • Quantified what would be needed: a modest number of additional samples processed in a design that breaks the confounding
  • Provided the sample size calculation and design for the follow-up at no additional charge
What the client received
  • A written report explaining the confounding, with the diagnostic figures that demonstrate it
  • A concrete recovery plan with sample numbers and processing design
  • An invoice for one week of work rather than a full analysis
  • Analysis of the extended cohort the following year, which did produce a publishable result

Timeline: 1 week to the finding; full analysis 4 weeks once the design was fixed

Genetic Variation

Variant Analysis in an African Family Cohort

A clinical genetics group investigating an undiagnosed paediatric condition across four consanguineous families, where standard pipelines returned an unmanageable candidate list.

What arrived

Whole exome sequencing on 14 individuals across four trios and extended families. An earlier analysis using default population frequency filters had returned several hundred candidate variants.

What we did
  • Rebuilt the filtering strategy using African-specific allele frequency resources rather than predominantly European reference panels
  • Applied inheritance models appropriate to consanguinity, including homozygosity mapping across the affected individuals
  • Ranked remaining candidates against HPO phenotype terms supplied by the clinical team
  • Flagged explicitly which candidate genes had thin annotation coverage in African populations, so absence of evidence was not read as evidence of absence
What the client received
  • A prioritised candidate list short enough to pursue experimentally
  • Runs-of-homozygosity plots supporting the inheritance model
  • An annotated VCF and a readable variant table with the evidence for each call
  • A methods section documenting every filter and threshold applied

Timeline: 5 weeks

Public Data Mining

Grant Preliminary Data From Public Repositories

A researcher needed preliminary data supporting a hypothesis for a grant deadline, with no budget to generate new sequencing.

What arrived

A gene of interest, a disease context, and eleven weeks until the submission deadline.

What we did
  • Searched GEO, ArrayExpress, SRA and TCGA against defined inclusion criteria, screening 60 candidate datasets down to nine usable ones
  • Reprocessed all nine from raw data through one pipeline so differences reflected biology rather than processing choices
  • Ran a meta-analysis across studies with study of origin modelled explicitly and heterogeneity quantified
  • Assembled an expression, mutation and survival profile for the gene across tissues and disease states
What the client received
  • Forest plots and a meta-analysis summary suitable for the grant application
  • A curated dataset inventory with accessions, ready to cite
  • Reprocessed matrices the group can reuse for follow-up questions
  • All processing code and the exact accession versions used

Timeline: 4 weeks, submitted with three weeks to spare

Single-Cell

Single-Cell Study: Fewer Clusters Than Expected

A group with a completed 10x experiment wanted cell type annotation and differential expression, and had already produced clusters they were preparing to publish.

What arrived

Six samples across two conditions, previously analysed with default parameters, yielding 22 clusters the group had begun annotating individually.

What we did
  • Reran quality control with ambient RNA correction and doublet detection, neither of which had been applied
  • Showed that six of the 22 clusters were driven by ambient contamination and doublets rather than distinct cell states
  • Reintegrated across donors with quantitative assessment that biological signal survived correction
  • Annotated the remaining clusters against reference atlases using both marker-based and reference-based methods
  • Ran differential expression as pseudobulk within cell type, rather than per cell, which had been inflating significance
What the client received
  • A defensible set of annotated clusters with the evidence for each call
  • Pseudobulk differential expression results per cell type
  • Documentation of every filtering and clustering decision, with the alternatives shown
  • An annotated object the group explores themselves

Timeline: 5 weeks

Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827