← All projects

Transcriptomics · Pipelines

RNA-seq pipeline with Nextflow and Docker

A samplesheet-driven Nextflow DSL2 pipeline that goes from paired-end FASTQ to differential expression in one command. Two quantification routes, a configurable DESeq2 design formula, every tool in a versioned Docker container, and a real-data validation run that recovers the textbook biology.

Context
Portfolio project
Data
Airway dataset GSE52778 (Himes et al. 2014): 4 donors, dexamethasone vs control, 8 libraries
Methods
FastQC, Cutadapt, STAR + featureCounts or Salmon, MultiQC, DESeq2 with paired design
Stack
Nextflow DSL2, Docker, R/DESeq2, GitHub Actions
Volcano plot of dexamethasone versus control in airway smooth muscle cells with known response genes labelled
Airway smooth muscle cells, dexamethasone vs control: 350 genes up and 265 down (padj < 0.05, |log2FC| > 1), known targets labelled.
8 / 8Canonical glucocorticoid-response genes recovered
615Differentially expressed genes (350 up, 265 down)
9.5 minSalmon-route run on a 16 GB laptop, with FastQC and Cutadapt cached
The problem

RNA-seq that runs the same way twice

RNA-seq analyses are often assembled by hand, depend on one machine’s environment and are hard to reproduce. I wanted a pipeline that runs end to end from a samplesheet, in containers, and proves itself on real data rather than just on a toy test.

Design

Two routes, one command

  1. QC and trimming. FastQC, then Cutadapt.
  2. Genome route. STAR alignment to sorted BAM, then fragment-level featureCounts.
  3. Transcript route. Salmon with gene-level summarisation, around 4 GB of RAM, so it fits a laptop.
  4. Differential expression. DESeq2 with a configurable design formula and an explicit reference level, so paired or blocked designs work and the sign of log2FC is unambiguous.
  5. Reporting. MultiQC across everything, plus MA and PCA plots.

Three versioned Docker images hold the tools. A tiny bundled test dataset runs in GitHub Actions on every push to check the wiring end to end.

Validation

Does it recover known biology?

I ran the pipeline on the public airway dataset: four donors’ airway smooth muscle cells, each untreated and treated with dexamethasone, with a paired ~ cell + condition design so each donor is compared with itself. Salmon mapped 93.9–95.3% of fragments in all eight samples.

Genelog2FCpadj
ZBTB167.369.6e−19
SPARCL14.642.7e−38
KLF154.446.5e−24
FKBP53.799.6e−21
TSC22D3 (GILZ)3.182.7e−10
PER13.171.2e−28
DUSP12.993.4e−81
CRISPLD22.693.0e−39

Every classic dexamethasone target is strongly induced, including CRISPLD2, the headline gene of the original study. On the PCA, treatment separates samples on PC1 (42% of variance); PC2 mainly separates one donor, which is exactly the effect the paired design removes.

PCA of the eight airway samples separating by treatment on PC1
PCA of the eight airway samples.
Notes

Limits worth knowing

  • The airway run used the first 3 million read pairs per library to stay laptop-sized; rerunning the download script with FULL=1 fetches the complete files.
  • The STAR route needs about 32 GB of RAM to build a human index (roughly 16 GB with a sparse suffix array). On a 16 GB laptop, use Salmon.

Get in touch

I’m looking for bioinformatics roles in statistical genetics, transcriptomics and NGS analysis, especially where clinical or cell and gene therapy experience helps. Based in London, open to hybrid and remote.