Nadja Eisenbruch DDSA PhD Fellow at the Technical University of Denmark

Nadja Eisenbruch

Position: Scaling Generative Enzyme Synthesis (GenES)
Categories: PhD Fellows 2026
Location: Technical University of Denmark

ABSTRACT:

Enzymes are central to the green transition, enabling sustainable manufacturing of chemicals and materials through highly selective catalysis under mild conditions and reducing reliance on energy-intensive processes. However, machine learning (ML)–driven enzyme design remains limited by insufficient functional data. Generative models can propose vast numbers of candidate sequences, yet experimental testing at gene length typically involves only ~10²–10³ individually synthesized variants, capturing only a small fraction of the distributions defined by modern protein models. Hence, how can we accurately translate generative protein models to experimental reality?
Variational synthesis overcomes this limitation by implementing generative models directly in DNA synthesis. Instead of synthesizing individual sequences, a synthesis model defines a distribution over protein sequences that is parameterized directly by experimentally controllable synthesis variables to closely approximate the target distribution p(x) defined by a trained generative model. Current implementations are restricted to short sequences (<70 amino acids), precluding application to full-length enzymes.
This PhD project will extend the published variational synthesis framework from oligonucleotides to kilobase-length genes by incorporating stochastic gene synthesis with PCR-based assembly. As a case study, we will focus on cytochrome P450-BM3 (CYP102A1), a widely used oxidative biocatalyst and model scaffold for enzyme engineering. We will train a synthesis-compatible generative model using evolutionary data and structure-guided design tools and apply sharpening strategies to bias the synthesis distribution toward functional regions of sequence space. The resulting synthesized library will be subjected to high-throughput functional screening and sequencing to collect activity and sequence data across ~10⁸ variants.

This data will be used to train a probabilistic sequence–activity model that accounts for experimental noise and enables quantitative inference and extrapolation across sequence space.

By integrating gene-length variational synthesis with large-scale functional learning, this project will generate the largest publicly available P450 BM3 sequence–activity dataset to date and establish a generalizable framework for translating generative protein models into experimental reality and thus accelerating data-driven enzyme engineering for sustainable biocatalysis.

DDSA