Public Data Mining & Meta-Analysis

The data that answers your question may already exist. We find it, reprocess it consistently, and analyse it — at a fraction of the cost of generating your own.

Millions of samples sit in public repositories, and a well-designed reanalysis can answer a research question, generate a hypothesis, or provide the preliminary data a grant application needs — without a single sample being collected.

The catch is that public data is inconsistently processed and unevenly annotated. Combining studies naively produces batch effects that swamp the biology. We reprocess everything through one pipeline and model study of origin explicitly.

Data We Accept

  • A research question — that is genuinely all we need to start
  • Any accession numbers you have already identified
  • Inclusion and exclusion criteria if you have them in mind
  • Your own data, if you want it placed in the context of public cohorts

Questions We Answer

  • Has anyone already generated the data I need?
  • Does my finding replicate across independent public cohorts?
  • How is my gene of interest expressed across tissues and diseases?
  • Can I get preliminary data for a grant without new sequencing?
  • What does the published evidence say when it is reanalysed consistently?

What We Do

Each project uses the subset of these that your research question requires.

01

Dataset Discovery and Curation

Systematic search of GEO, ArrayExpress, SRA, ENA, dbGaP, TCGA, GTEx, ICGC, the GWAS Catalog and African resources including H3Africa, with screening against your criteria and a documented selection trail.

02

Uniform Reprocessing

Every dataset run through the same pipeline from raw data where available, so differences between studies reflect biology rather than processing choices.

03

Cross-Study Integration

Batch correction and meta-analysis across studies, with study of origin modelled explicitly and heterogeneity quantified rather than ignored.

04

Target and Biomarker Landscaping

Expression, mutation and survival profiles for a gene or gene set across tissues, cancers and disease states, assembled into a single evidence summary.

05

Benchmarking Your Own Data

Placing your cohort alongside public reference cohorts to check whether your samples behave as expected, and whether your finding replicates.

What You Receive

  • Curated dataset inventory with accessions, sample counts and metadata
  • Uniformly reprocessed expression, variant or summary matrices
  • Meta-analysis results with heterogeneity statistics and forest plots
  • Evidence summary report for your gene, variant or pathway of interest
  • All processing code and the exact accession versions used

Tools We Use

  • GEOquery, SRA Toolkit, fasterq-dump, pysradb
  • recount3, ARCHS4, refine.bio
  • TCGAbiolinks, cBioPortal, GTEx
  • metafor, MetaVolcanoR
  • ComBat, RUVSeq, Harmony

Typical turnaround: 2–5 weeks depending on the number of datasets included

Indicative price: From $599 for a single-gene landscaping report

Reduced rates are available for students and researchers at African public institutions. Every project is quoted in writing before work begins.

Services are provided for research purposes only. They are not intended for clinical diagnosis, treatment decisions or individual health assessment. See how it works, data submission guidelines and what you receive.

Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827