The program asks how machine learning can narrow biological search spaces without losing sight of the experiment that must eventually resolve uncertainty.
Flagship problem: oxygen-tolerant nitrogen fixation
Nitrogen fixation is an unusually clear test of this approach. Nitrogenase is oxygen-sensitive, while cyanobacteria produce oxygen through photosynthesis. Understanding how some organisms coordinate those processes requires functional discovery across genes, regulatory context, environmental condition, and phenotype.
Oxygen-tolerant nitrogen fixation is the flagship application. The same framework can support other problems where biological function is undercharacterized and experimental capacity is constrained.
A connected methodological program
Functional discovery
Begin with a biological function and the evidence required to distinguish a promising candidate from a convenient correlation.
Environmental and condition-specific biology
Expression and phenotype depend on context. Condition-specific transcriptomic and proteomic evidence helps define when a mechanism is relevant.
Sequence and protein representation
Comparative genomics and protein-language-model representations make it possible to reason across related genes and proteins even when annotation is incomplete.
Multi-omic integration
No single data layer settles the question. The task is to combine partially informative evidence while keeping its provenance visible.
Experimental prioritization
The output is a smaller, defensible set of candidates and a clear account of why each candidate is worth testing.
Iterative model-and-experiment cycles
Active learning treats each experimental result as the beginning of the next round, building a more efficient learning system over time.
Protein learning and engineering
For protein-property questions, the same logic moves from sequence representation through candidate selection to validation planning.