← All projects

Regulatory genomics · Machine learning

RegulonML: predicting regulatory variant effects in unseen promoters and enhancers

Can we predict how a single-base change in a promoter or enhancer alters activity, for a regulatory element that wasn’t in the training data? Using a saturation-mutagenesis MPRA, the model sees only DNA sequence and is evaluated with leave-one-locus-out cross-validation, so every prediction is for DNA it has never seen.

Context
Portfolio project
Data
Kircher et al. 2019 saturation mutagenesis MPRA: 21 elements, 41,724 variants after QC
Methods
Motif gain/loss scoring (1,019 JASPAR 2026 motifs), ridge, gradient boosting, leave-one-locus-out CV
Stack
Python, scikit-learn, pyjaspar, pytest
Held-out performance by model; the hatched bar is the leaky random split.
Held-out performance by model; the hatched bar is the leaky random split.
0.07 → 0.13Median held-out Spearman ρ for gradient boosting when motif features are added
0.60AUROC for significant variants on unseen loci
0.31The inflated ρ a leaky random split gives, shown for comparison
The question

Generalising to regulatory DNA the model hasn’t seen

Kircher et al. measured every possible single-nucleotide substitution and 1-bp deletion in 21 disease-associated regulatory elements. The hard, useful question isn’t fitting those elements; it’s predicting variant effects in a new element from sequence alone.

Design

Built so test data can’t leak into training

MPRA effects are calculated from DNA and RNA read counts, and the same element appears in several datasets (TERT in four cell lines, SORT1 in three), so leakage is easy to introduce. Every part of the pipeline is set up to prevent it.

  • Sequence-only features. Read counts are used only for QC filtering, and a test enforces this.
  • Grouped by locus. All datasets from the same element are held out together, so no DNA appears in both training and test.
  • No genome download. Each element’s reference sequence is rebuilt from the mutagenesis data itself.
  • Motif scoring. Binding-site creation and loss for every SNV and deletion against all JASPAR 2026 vertebrate motifs, on both strands, with a vectorised scorer checked against brute force.
  • Explicit baselines. Per-element Spearman, AUROC and AUPRC against a mean baseline, with a random split reported alongside to show what leakage would do to the numbers.
Results

Modest, real gains on unseen DNA

FeaturesModelMedian ρAUROCAUPRC
—Mean baseline—0.500.18
Sequence contextGradient boosting0.070.580.21
Context + TF motifsRidge0.140.550.22
Context + TF motifsGradient boosting0.130.600.24
Context + TF motifs (random split, leaky)Gradient boosting0.31——

Motif features roughly double the gradient-boosting model’s held-out correlation and improve 15 of 21 loci. The best models rank variant effects better than chance in 28 of 29 held-out datasets. Letting the same element appear in training and test more than doubles ρ to 0.31, which shows how far a random split would overstate performance.

Bar chart of held-out Spearman correlation for each locus
Held-out Spearman ρ per locus.
Biology check

The TERT promoter, with no training

The two recurrent cancer mutations in the TERT promoter, C228T and C250T, are the strongest activating variants in the assay and are known to create new ETS-family binding sites. The motif scoring recovers this with no training: among 878 TERT variants they rank 3rd and 11th (joint) for ETS-site creation. With TERT held out, the model ranks C228T 15th of 878 for activation but misses C250T, because activating variants are rare in the training loci.

Two-panel plot of the TERT promoter: measured MPRA effect and ETS-family site creation by position, with C228T and C250T highlighted
Saturation mutagenesis of the TERT promoter: measured effect (top) and ETS-family site creation (bottom), C228T/C250T highlighted.
Next steps

What would improve it

  • Motif scores ignore which transcription factors are expressed in each cell line; cell-type TF expression is the obvious next feature.
  • A deep sequence model (for example Enformer or Borzoi ref/alt predictions) is the natural comparison.
  • Only 21 loci: per-locus results are noisy and the model rarely sees activating variants.

Get in touch

I’m looking for bioinformatics roles in statistical genetics, transcriptomics and NGS analysis, especially where clinical or cell and gene therapy experience helps. Based in London, open to hybrid and remote.