← All projects

Single-nucleus transcriptomics · Parkinson’s disease

GPNMB in Parkinson’s disease: single-nucleus RNA-seq re-analysis

A focused re-analysis of human substantia nigra pars compacta single-nucleus RNA-seq to find where GPNMB is expressed, whether its microglial expression differs between Parkinson’s disease and controls, and whether independent published analyses point the same way. It adds cell-type context to the genetic evidence from my dissertation.

Context
Independent follow-up to my MSc dissertation
Data
GSE243639 (Martirosyan et al.): 29 post-mortem SNpc donors, 15 PD and 14 controls
Methods
Donor-level pseudobulk DE, microglial re-clustering, propeller composition, enrichment, external validation
Stack
Python, Scanpy, AnnData, PyDESeq2; R speckle
UMAP of 83,484 substantia nigra nuclei coloured by author cell-type annotation
GSE243639 author cell-type annotations, 83,484 nuclei.
83,484Nuclei across 29 donors (author-QC atlas, 33,538 genes)
~2×Higher GPNMB in PD microglial pseudobulk (log2FC 1.03, P = 0.009)
0.22Transcriptome-wide FDR for GPNMB: nominal, not significant after correction
The question

Where is GPNMB expressed, and does it change in PD?

Genetic evidence can prioritise a protein, but it doesn’t say which cells matter. I asked in which substantia nigra cell populations GPNMB is expressed, and how its expression and the associated microglial programme differ between Parkinson’s disease and controls. GPNMB was a pre-specified target, not a hit picked from transcriptome-wide results.

Principles

Donors, not nuclei, are the replicates

Cell-level differential expression can overstate evidence when thousands of nuclei from the same donor are treated as independent. The analysis rules were fixed before looking at disease effects and stored in config.yaml so changes stay auditable.

  • Formal PD–control tests use donor-level pseudobulk counts from raw integers; cell-level plots are descriptive only.
  • At least 30 nuclei per donor and at least 5 eligible donors per condition before testing.
  • Primary design ~ age + sex + condition, deliberately not overfitting a 29-donor cohort.
  • Harmony off by default: it can help embeddings but can also erase real donor biology, and never feeds DE counts.
  • External validation kept separate from discovery.
Localisation

Highest in microglia, with OPCs close behind

UMAP of all nuclei coloured by GPNMB expression
GPNMB expression across the atlas (log1p normalised).
Box plot of mean GPNMB expression in microglia per donor, control versus PD
Donor-level mean GPNMB in microglia, control vs PD. Each point is one donor.
Cell typeDonorsMedian donor expressionDetection fractionNuclei
Microglia290.28900.184412,995
OPC290.17320.17176,644
Vascular290.08010.06201,560
Astrocytes290.02330.016920,710
Neurons290.02130.02696,196
Oligodendrocytes290.00820.007535,032
T cells260.00000.0000347
PD vs control

About two-fold higher in PD microglia, nominally

Donor-level pseudobulk on 14 control and 15 PD donors, adjusted for age and sex, gave log2FC +1.025 (SE 0.390; P = 0.00863), roughly 2.03-fold higher GPNMB in PD microglia. It does not survive transcriptome-wide FDR (0.218).

Broad microglial abundance was not detectably different between groups (propeller P = 0.587), so the higher signal isn’t simply more microglia in PD donors. Neuronal proportion was lower in PD (P = 0.009) but did not cross FDR < 0.05 (0.062).

Six UMAP panels of 12,995 re-clustered microglial nuclei: Leiden clusters, condition, GPNMB and three programme scores
12,995 microglial nuclei re-clustered de novo: Leiden clusters, condition, GPNMB, and homeostatic, GPNMB-associated activation and interferon programme scores.
Power

Underpowered for transcriptome-wide significance, not evidence against an effect

A nominal P with FDR 0.218 can mean a small effect or a small cohort. To tell them apart I simulated donor-level GPNMB pseudobulk counts from a negative binomial calibrated to this result: the dispersion was back-solved so the real ~ age + sex + condition design reproduces the observed standard error (0.390), and age and sex were resampled from the real donors. Each simulated cohort was re-tested with a Wald test, 5,000 times per scenario. Under no effect the false-positive rate stayed between 4.8% and 6.1%.

Two bars matter. As the pre-specified target, GPNMB only needs P < 0.05. To survive transcriptome-wide correction it needed P ≤ 4.9 × 10−4, the Benjamini–Hochberg cut-off this analysis actually used (141 of 14,358 genes at FDR 0.05).

Two panels of simulated power against donors per group for true fold changes of 1.4, 1.7 and 2. At P < 0.05, a two-fold effect reaches 80% power at 20 donors per group; at the transcriptome-wide FDR cut-off it needs 50.
Simulated power by donors per group. Dotted vertical line: this study (14 vs 15 donors).
True effectPower now, P < 0.05Power now, FDRDonors/group for 80%, P < 0.05Donors/group for 80%, FDR
1.4× (log2FC 0.5)25%2%70>100
1.7× (log2FC 0.75)48%6%3580
2.0× (log2FC 1.0, as observed)71%18%2050

Even if the observed two-fold effect is real, a 29-donor cohort had about an 18% chance of clearing transcriptome-wide FDR, so the non-significant FDR is what this design would usually produce. Confirming it at FDR would take roughly 50 donors per group. This bar is conservative, because a larger cohort would reject more genes and loosen the cut-off. The dispersion is held at its calibrated value rather than re-estimated in each simulation, which makes the estimates slightly optimistic.

External validation

Same direction, variable strength

DatasetMethodlog2FCPFDR
GSE243639 (this analysis)Donor-level pseudobulk DESeq2+1.0250.008630.218
GSE184950 / WangDonor-aware microglial MAST+0.1130.1311.000
Wang + Kamath + SmajicMulti-dataset microglial pseudobulk+0.9330.001860.0501

GSE184950 is used through its published supplementary tables, because the public Seurat object lacks the final cell-type labels needed for a defensible microglial pseudobulk.

What it supports

  • GPNMB expression concentrated in microglia.
  • A positive, nominally significant PD effect in donor-level pseudobulk.
  • Directional consistency across independent studies.
  • That the missed FDR is expected at 29 donors (≈18% power for a two-fold effect).

What it doesn’t show

  • That GPNMB causes PD, or whether higher GPNMB is harmful or protective.
  • A significant replication in GSE184950 alone.
  • That a specific microglial state mediates the genetic association.

Get in touch

I’m looking for bioinformatics roles in statistical genetics, transcriptomics and NGS analysis, especially where clinical or cell and gene therapy experience helps. Based in London, open to hybrid and remote.