DataCore Analytics

nf-core Pipeline Reference

Sixteen nf-core pipelines, what each one is for, and an honest note on whether it will run on a laptop.

nf-core is a community of researchers who build and peer-review reproducible bioinformatics workflows. The pipelines are free, versioned, continuously tested and citable. They are, in our view, the single most useful thing that has happened to reproducible bioinformatics in the last decade.

The table below is our working reference: what each pipeline does, what data it takes, and the practical constraint that usually bites first. Every samplesheet schema referenced here was verified against the pipeline's own repository.

Our Analysis Workbench will build a validated samplesheet and generate the exact run command for any of these, to run on your own machine.

The Catalogue

RAM shows the reduced configuration through to the full one. Disk is headroom for references, intermediates and the Nextflow work directory.

PipelineAreaWhat it doesRAMDisk
Bulk RNA-Seq
nf-core/rnaseq
TranscriptomicsQuality control, trimming, alignment and quantification of bulk RNA sequencing, producing a gene expression matrix and an extensive QC report.16–38 GB120 GB
Differential Expression
nf-core/differentialabundance
TranscriptomicsTakes a count matrix and a sample sheet and runs differential analysis with DESeq2 or limma, producing volcano plots, heatmaps, enrichment and a shareable HTML report.4–8 GB5 GB
Single-Cell RNA-Seq
nf-core/scrnaseq
Single-cellPre-processing of droplet and plate-based single-cell RNA-Seq into a cell-by-gene count matrix, with Alevin, Kallisto|Bustools, STARsolo or CellRanger.16–32 GB150 GB
Variant Calling (WGS / WES)
nf-core/sarek
GenomicsGermline and somatic variant calling from whole genome, exome or targeted sequencing, following GATK best practices with annotation by VEP or snpEff.16–64 GB500 GB
Rare Disease (Trios)
nf-core/raredisease
GenomicsCall, annotate and rank variants from WGS or WES for rare disease diagnostics, with trio-aware filtering and phenotype-driven prioritisation.16–64 GB400 GB
Viral Genome Surveillance
nf-core/viralrecon
Pathogen genomicsAssemble consensus viral genomes from amplicon or metagenomic sequencing, call variants and assign lineages — the standard route for SARS-CoV-2 surveillance.8–16 GB50 GB
Bacterial Assembly
nf-core/bacass
Pathogen genomicsAssemble, annotate and type bacterial genomes from short reads, long reads or a hybrid of both.8–32 GB60 GB
16S / ITS Amplicon
nf-core/ampliseq
MicrobiomeAmplicon sequencing analysis with DADA2 and QIIME2: denoising to exact sequence variants, taxonomic assignment and diversity analysis.8–16 GB30 GB
Shotgun Taxonomic Profiling
nf-core/taxprofiler
MicrobiomeTaxonomic classification and profiling of shotgun metagenomes across multiple classifiers at once, with standardised output.16–100 GB200 GB
Metagenome Assembly & Binning
nf-core/mag
MicrobiomeAssembly, binning and annotation of metagenomes to recover metagenome-assembled genomes, from short reads, long reads or both.32–128 GB400 GB
DNA Methylation (Bisulfite)
nf-core/methylseq
EpigenomicsBisulfite sequencing analysis with Bismark or bwa-meth, producing per-cytosine methylation calls and coverage reports.16–32 GB200 GB
ChIP-Seq
nf-core/chipseq
EpigenomicsChIP-seq peak calling and QC: alignment, filtering, MACS peak calling against matched controls, and differential binding analysis.16–32 GB150 GB
ATAC-Seq
nf-core/atacseq
EpigenomicsChromatin accessibility analysis: alignment, mitochondrial and duplicate filtering, peak calling and differential accessibility.16–32 GB150 GB
CUT&RUN / CUT&Tag
nf-core/cutandrun
EpigenomicsAnalysis of CUT&RUN and CUT&Tag data, including spike-in normalisation and SEACR or MACS2 peak calling.16–32 GB120 GB
AMR & Biosynthetic Gene Screening
nf-core/funcscan
Pathogen genomicsScreen assembled contigs for antimicrobial resistance genes, antimicrobial peptides and biosynthetic gene clusters.8–24 GB40 GB
Fetch Public Data
nf-core/fetchngs
Data acquisitionDownload raw FASTQ files and metadata from SRA, ENA, DDBJ or GEO by accession, and auto-generate a samplesheet for the next pipeline.4–8 GB500 GB

Running These On A Laptop: Transcriptomics

The constraint that usually bites first.

01

Bulk RNA-Seq — nf-core/rnaseq

Aligning to a human genome with STAR needs roughly 38 GB of RAM to build the index. On a laptop use --pseudo_aligner salmon with --skip_alignment, or a smaller genome.

02

Differential Expression — nf-core/differentialabundance

Runs comfortably on any modern laptop — it works from a count matrix, not raw reads.

Running These On A Laptop: Single-cell

The constraint that usually bites first.

01

Single-Cell RNA-Seq — nf-core/scrnaseq

Alevin and Kallisto are far lighter than CellRanger. Expect long runtimes on a laptop for more than a couple of samples.

Running These On A Laptop: Genomics

The constraint that usually bites first.

01

Variant Calling (WGS / WES) — nf-core/sarek

Whole genome sequencing is not realistic on a laptop. Exomes with a target BED are feasible overnight; anything larger needs a server or cloud.

02

Rare Disease (Trios) — nf-core/raredisease

Reference-heavy. Plan for a server unless you are running a single exome.

Running These On A Laptop: Pathogen genomics

The constraint that usually bites first.

01

Viral Genome Surveillance — nf-core/viralrecon

Genuinely laptop-friendly. Viral genomes are small — a full plate of 96 amplicon samples is realistic overnight.

02

Bacterial Assembly — nf-core/bacass

Feasible for a handful of isolates. Skip Kraken2 unless you have the database locally — it alone wants over 60 GB.

03

AMR & Biosynthetic Gene Screening — nf-core/funcscan

Works from assemblies rather than reads, so it is comparatively light.

Running These On A Laptop: Microbiome

The constraint that usually bites first.

01

16S / ITS Amplicon — nf-core/ampliseq

One of the most laptop-friendly pipelines here. A 96-sample 16S run is comfortable overnight.

02

Shotgun Taxonomic Profiling — nf-core/taxprofiler

The constraint is the database, not the pipeline. A standard Kraken2 database needs over 60 GB of RAM; use a capped or reduced database on a laptop.

03

Metagenome Assembly & Binning — nf-core/mag

The heaviest pipeline in this catalogue. Metagenome assembly is memory-hungry — plan for a server unless your samples are very small.

Running These On A Laptop: Epigenomics

The constraint that usually bites first.

01

DNA Methylation (Bisulfite) — nf-core/methylseq

Building a Bismark index for a human genome is slow and memory-hungry. Build it once, keep it, and pass it with --bismark_index.

02

ChIP-Seq — nf-core/chipseq

Feasible for a small number of samples if the aligner index already exists.

03

ATAC-Seq — nf-core/atacseq

Similar footprint to ChIP-seq. Pre-build the aligner index.

04

CUT&RUN / CUT&Tag — nf-core/cutandrun

CUT&RUN libraries are usually shallow, which makes this lighter than ChIP-seq.

Running These On A Laptop: Data acquisition

The constraint that usually bites first.

01

Fetch Public Data — nf-core/fetchngs

Limited by your bandwidth and disk, not your processor. Check the total size on ENA before starting.

Practical Notes That Apply To All Of Them

Cap your resources or tasks will fail

nf-core defaults assume a compute cluster. Without a config capping CPUs and memory to what you actually have, tasks request more than the machine offers and die. The workbench generates that config for you.

The work directory grows fast

Nextflow writes every intermediate into work/. A run that produces 2 GB of results can leave 80 GB behind. Keep it until you are finished — it is what makes -resume work — then delete it.

Build reference indices once

Aligner indices for a human genome take hours and a lot of memory to build. Build once, keep them, and pass the path explicitly on subsequent runs.

Use -resume, always

It costs nothing and means an interrupted run picks up from the last completed step rather than starting again. On a laptop that will close its lid at some point, this matters.

Containers beat Conda

Use -profile docker or singularity where you can. Conda environments resolve slowly and are a common source of irreproducibility between machines.

Test with -profile test first

Every nf-core pipeline ships a tiny test dataset. Running it confirms your installation works before you commit days of compute to real data.

When To Stop And Ask

Running a pipeline is the easy part. Consider getting help when:

  • You are not sure the study design can answer your question — this is worth asking before data generation, not after
  • Your samples do not cluster the way the design predicts, and you do not know whether it is batch or biology
  • You are working with African population data and using reference panels or prediction models built elsewhere
  • The pipeline completed but you cannot tell whether the result is real
  • A reviewer has challenged your analysis and you need an independent opinion

The scoping call is free, and we will tell you plainly if you do not need us.

Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827

Free scoping call · reply within 1 working day Get a Quote