Mercor
Bioinformatician · Scientific AI Evaluation
Designed scientific-reasoning evaluations that test whether frontier language models can infer hidden biological and computational rules through controlled experimentation.
- Authored 106 accepted tasks (Jan–Jul 2026) for a frontier model’s bioinformatics benchmark. Each is a hidden Python function implementing a cited genomics or biostatistics rule, with unit tests, that the model must reverse-engineer by choosing inputs and reading outputs.
- Calibrated difficulty from model trajectory data using ablation ladders and adversarial foils: textbook solutions scored below 50% while the full solution reached 100%.
- 80% first-pass approval through peer and client review.
- Expert evaluation of AI agent runs on preclinical drug R&D tasks, and blinded A/B comparisons of coding models on large open-source codebases.