Data Submission Guidelines

How to send us your data: accepted formats, transfer methods, the metadata we need, and what we check on arrival.

We are an analysis-only organisation and never handle physical samples. Everything below concerns digital data transfer.

Read this before your first transfer. Most delays at the start of a project are caused by missing metadata or incomplete file sets, not by the analysis itself.

The Transfer Process

Data is only accepted after an agreement is signed.

01

Agreement First

No data is accepted before a written agreement is in place covering scope, storage location, retention period, access, authorship and intellectual property. Where the study involves human subjects we ask to see ethical approval and any applicable data transfer agreement.

02

You Receive a Transfer Link

We send an encrypted, time-limited upload link to a Google Cloud bucket provisioned for your project alone. Do not email data, and do not send it through consumer file-sharing services.

03

Upload Data and Metadata Together

Upload your files along with the completed metadata sheet we send you. Incomplete metadata is the single most common cause of delay.

04

We Verify on Arrival

Within two working days we confirm receipt, verify checksums, check that every expected file is present and readable, and run initial quality control. If something is missing or corrupted we tell you immediately.

05

Analysis Begins

Once verification passes, your named analyst starts work to the approved plan and confirms the delivery date.

06

Retention and Deletion

At the end of the project data is returned, retained for the agreed period, or securely deleted — whichever your agreement specifies. We confirm deletion in writing.

Accepted File Formats

Data typeFormats we acceptNotes
Sequence readsFASTQ (.fastq.gz, .fq.gz) — gzip compressedKeep read pairs as separate R1/R2 files with consistent naming
Aligned readsBAM, CRAMInclude the index (.bai / .crai) and tell us the reference genome build used
Nanopore signalPOD5, FAST5Only needed if you want rebasecalling; otherwise send FASTQ
VariantsVCF, gVCF, BCF (bgzip compressed)Include the .tbi index and the reference build
Expression matricesCSV, TSV, MTX, RDS, H5ADState whether values are raw counts, TPM, FPKM or normalised
Single-cellCellRanger outs directory, H5AD, RDS, LoomSend the filtered and raw matrices where both exist
Methylation arraysIDAT (both Grn and Red per sample)Include the sample sheet with sentrix ID and position
Genotype arraysPLINK (.bed/.bim/.fam), VCF, Illumina final reportState the array and genome build
Clinical / tabularCSV, XLSX, Stata (.dta), SPSS (.sav), REDCap exportSend the data dictionary with it
Reference filesFASTA, GTF, GFF3Only if you need a non-standard reference or annotation

Metadata We Need

The sample sheet

Every transfer must include a sample sheet as CSV or XLSX with one row per sample. Without it we cannot begin, because file names alone rarely tell us which sample belongs to which group.

  • Sample ID — exactly matching the file names, no spaces or special characters
  • File names — the exact files belonging to that sample, including R1 and R2
  • Group or condition — the comparison you want to make
  • Batch information — sequencing run, extraction date, plate, chip position
  • Covariates — age, sex, site, timepoint, or anything else that may confound
  • Collection date and location — required for surveillance and phylogenetic work

Study information

Alongside the sample sheet, tell us:

  • The organism and reference genome build you expect us to use
  • The library preparation kit and sequencing platform
  • Whether the data is stranded, paired-end, and the read length
  • Any samples you already know are problematic, and why
  • The research question, restated in one or two sentences

Naming conventions

Use only letters, numbers, hyphens and underscores in file names. Avoid spaces, brackets, ampersands and non-ASCII characters — they break pipelines silently. Keep the sample identifier at the start of the file name and consistent across all files belonging to that sample.

Checksums

Generate MD5 or SHA-256 checksums before upload and send them with the data. Large transfers do occasionally corrupt, and a checksum mismatch caught on day one is far cheaper than a result you cannot reproduce three weeks later.

Data Security and Human Subjects

Where your data is stored

Data is held on Google Cloud with encryption in transit and at rest, in a project-specific bucket. Access is limited to your named analyst and the reviewing consultant, and every access is logged. We do not store client research data on personal devices.

De-identification

Please de-identify data before transfer wherever the analysis allows it. Replace names, hospital numbers and national identifiers with study codes, and keep the linking key at your own institution — we do not need it and prefer not to hold it.

Where identifiable data is genuinely unavoidable for the analysis, the handling terms are set out explicitly in the project agreement before transfer.

Ethical approval

For projects involving human subjects we ask for the approval reference and approving committee, and a copy of the approval letter. This is not bureaucracy for its own sake: journals ask for it, and it is easier to produce at the start than during revision.

What we will not do

We do not reuse client data for other projects, for method development or for training machine learning models without written permission. We do not share data between clients. We do not deposit data in public repositories on your behalf unless you ask us to and the consent framework allows it.

Common Problems and How to Avoid Them

These are the issues that most often delay a project by a week or more:

  • Missing R2 files. Paired-end datasets arriving with only forward reads — check file counts before uploading
  • Sample sheet does not match file names. Trailing spaces and inconsistent capitalisation are the usual culprits
  • Unknown genome build. Data aligned to an unstated reference cannot be safely combined with anything else
  • Normalised values sent as raw counts. Differential expression tools require raw counts; TPM will produce wrong results silently
  • No batch information. If we cannot model batch, we cannot rule it out as the explanation for your finding
  • Zipped archives of zipped archives. Send files as they are; nested archives slow verification and sometimes corrupt
Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827