Gabriel Sainz Vazquez DDSA PhD Fellow at the University of Copenhagen

Gabriel Sainz Vazquez

Position: Geometry-aware generative priors for sample-efficient Bayesian optimization in protein design
Categories: PhD Fellows 2026
Location: University of Copenhagen

ABSTRACT:

Protein and molecular design can enable new therapeutics, enzymes, and materials, but progress is often limited by a practical bottleneck: experimental evaluation is expensive, noisy, and slow. Wet-lab assays and high-fidelity computational proxies can test only a small number of candidates per round, so discovery depends on choosing the right few variants to measure and learning efficiently from noisy feedback.
Many iterative design pipelines combine Bayesian optimization (BO) (a data-efficient way to choose what to test next from past measurements) with a generative model that proposes new valid sequences. The key assumption is that “similar sequences behave similarly,” so results from tested variants can guide decisions about new ones. In practice, this is fragile in discrete sequence design: a single mutation can change function dramatically, and generic notions of similarity can lead to confident mistakes exactly where they matter: when picking the next experiments. In addition, pushing the generator too aggressively toward promising, high-expected-performance candidates can produce many look-alike candidates, and candidate generation itself can become too slow or compute-heavy to run many rounds.
This project aims to address these pitfalls in three practical steps: (1) use the generator to define which variants are “neighbors,” so we generalize along realistic mutation steps; (2) keep each proposed batch varied with diversity control, and adjust how strongly we bias proposals toward high predicted performance based on straightforward reliability checks from recent rounds; and (3) keep the loop practical by comparing diffusion- and flow-based sampling under matched experiment budgets and matched GPU-hour budgets.
We will validate on open-source closed-loop benchmarks (poli/poli-baselines), reporting best-found value and hit-rate alongside reliability on selected candidates, diversity/duplication, constraint feasibility, and compute cost, with ablations to attribute gains. We will also stress-test the approach in real-world protein engineering case studies (enzyme optimization) with the host experimental environment (BRIGHT, DTU) to evaluate robustness under realistic assay noise and practical constraints. Code, benchmark runs, and logs will be released openly, and experimental data will be shared as permitted.

DDSA