Plain-language definitions of the terms that come up most often in our scoping calls.
Written for researchers who are strong in their own field and newer to genomic data analysis. Where a term is commonly misused, we say so rather than just defining it.
If something you need is missing, tell us and we will add it.
Adjusted p-value (FDR)
The p-value corrected for the fact that you tested thousands of features at once. With 20,000 genes and no correction, roughly 1,000 will look significant at p < 0.05 by chance alone. Always report adjusted values; Benjamini-Hochberg is the usual method.
Alignment rate
The percentage of sequencing reads that map to the reference genome. Low rates usually mean contamination, the wrong reference, or poor library quality — investigate before interpreting anything downstream.
Allele frequency
How common a variant is in a population. The single most important filter in rare disease analysis, and the one most often applied wrongly to African samples using European reference panels.
Amplicon sequencing
Sequencing a specific PCR-amplified region rather than everything present. 16S for bacteria and ITS for fungi are the common examples, as is ARTIC tiling for viral genomes.
ASV (Amplicon Sequence Variant)
An exact sequence recovered from amplicon data after denoising, replacing the older OTU concept. ASVs are comparable between studies in a way that OTUs clustered at 97% identity are not.
Batch effect
Systematic technical variation between groups of samples processed together. If batch is confounded with your experimental condition, the study cannot distinguish them and no analysis will rescue it. Design it out at the randomisation stage.
Calibration
Whether a model's predicted probabilities match observed frequencies. A model can rank patients perfectly and still be badly calibrated, which makes it unsafe for decisions about individuals.
Coverage / depth
How many times each base is sequenced on average. Roughly 30x is standard for germline whole genome calling; somatic variant detection in tumours needs far more.
CPM / TPM / FPKM
Normalised expression units correcting for library size and, for TPM and FPKM, gene length. Use raw counts for differential expression tools — they model the count distribution themselves. Never feed TPM to DESeq2.
De novo assembly
Reconstructing genomes from reads without a reference. Memory-hungry, especially for metagenomes, and the usual reason a laptop run fails.
Differential expression
Testing which genes change between conditions. The statistics matter: negative binomial models such as DESeq2 or edgeR, not a t-test on normalised values.
FASTQ
The standard format for raw sequencing reads: sequence plus a per-base quality score. Almost always gzipped, and almost always the starting point.
FRiP score
Fraction of Reads in Peaks — the key quality metric for ChIP-seq and ATAC-seq. A low FRiP means your enrichment did not work, regardless of how many peaks were called.
GWAS
Genome-wide association study: testing millions of variants for association with a trait. Requires careful control of population structure, which is especially important in admixed African cohorts.
Imputation
Inferring genotypes not directly measured, using a reference panel. Accuracy depends heavily on the panel matching your population — a persistent problem for African cohorts.
Lineage assignment
Classifying a pathogen genome into a named lineage, such as Pango for SARS-CoV-2. Depends on curated definitions that are usually built from globally unrepresentative sequence sets.
MultiQC
A tool that aggregates quality metrics from an entire pipeline into one report. The first thing to open when a run finishes, and the file our visualiser reads.
Nextflow
The workflow manager nf-core pipelines are written in. Handles parallelism, containers, resuming and provenance so you do not have to.
nf-core
A community that builds peer-reviewed, versioned, continuously tested bioinformatics pipelines. Free, citable, and the basis of our workbench.
PCA (Principal Component Analysis)
A method for reducing many measurements to a few axes capturing the most variation. In practice, the first plot to look at: if replicates do not group together, stop and find out why.
Pseudoalignment
Quantifying transcript abundance without full genome alignment, as Salmon and Kallisto do. Far lighter on memory, which is what makes RNA-Seq feasible on a laptop.
Polygenic risk score
A score combining many variants to estimate trait risk. Loses most of its predictive power when transferred across ancestry groups, which is why scores built on European cohorts should not be applied to African populations without local validation.
Reference bias
The systematic failure of analyses to represent samples that differ from the reference genome. Because the human reference poorly represents African haplotypes, reads from divergent regions align badly or not at all.
Reproducibility
Whether someone else — including you in two years — can rerun your analysis and get the same numbers. Requires the code, the software versions, the reference versions and the parameters, not just the results.
Samplesheet
The CSV that tells an nf-core pipeline which files belong to which sample and how they are grouped. The most common cause of a run failing in the first minute.
Single-cell RNA-Seq
Measuring gene expression in individual cells rather than averaged across a tissue. Powerful, expensive, and unusually easy to over-interpret — cluster count is a parameter choice, not a discovery.
Strandedness
Whether an RNA library preserves which DNA strand a transcript came from. Getting it wrong silently halves your counts. Set it to 'auto' unless you are certain.
Structural variant
A large genomic change — deletion, duplication, inversion or translocation — as opposed to a single base change. Harder to call reliably, and often what is found when exome sequencing returns nothing.
Variant calling
Identifying differences between a sample and the reference genome. Germline calling finds inherited variants; somatic calling finds those acquired in tissue such as a tumour.
VCF
Variant Call Format: the standard file for storing called variants, usually bgzipped with a .tbi index alongside it.
Volcano plot
Effect size against statistical significance, one point per feature. Useful at a glance, but a wide, symmetrical cloud with nothing significant often means an underpowered study rather than no biology.
Whole exome sequencing
Sequencing only the protein-coding regions, roughly 1–2% of the genome. Much cheaper than whole genome sequencing and sufficient for most rare disease diagnostics.
