ForetoData / research

Choosing better experiments in complex biological systems.

Represent → Simulate → Decide → Learn

Biological discovery often means making expensive experimental choices with incomplete evidence. I use scientific machine learning to represent what is known, compare possible experiments, and decide which next test is most likely to improve understanding.

The goal is not simply to rank biological candidates. It is to build a decision system that connects computational evidence to a sequence of experiments that steadily reduces uncertainty and improves the next choice.

Represent the biological system

Nitrogen fixation is an unusually clear test of this approach. Nitrogenase is oxygen-sensitive, while cyanobacteria produce oxygen through photosynthesis. Understanding how some organisms coordinate those processes requires functional discovery across genes, regulatory context, environmental condition, and phenotype.

I represent the biological function, the conditions in which it matters, the available measurements, and the gaps that experiments must resolve. Oxygen-tolerant nitrogen fixation is the flagship application, but the same structure applies wherever biological function is undercharacterized and testing capacity is limited.

Simulate and compare possibilities

No single data source settles the question. Scientific machine learning combines comparative genomics, condition-specific transcriptomics and proteomics, protein-language-model representations, and other biological evidence to estimate what different candidate choices could reveal while keeping the origin and limits of each signal visible.

Decide what to test next

The practical output is a smaller, defensible set of candidates and a clear reason to test each one. Optimization helps balance predicted value, uncertainty, experimental cost, and the information a result could add—not just select the highest model score.

Learn from every result

Experimental evidence updates the model of the system and changes what should be tested next. Active learning turns that feedback into a continuing cycle, whether the question involves gene function, regulatory context, or protein properties and engineering.

Figure 01 / choosing the next experiment
The goal is not only a prediction. It is a better, testable next experiment.

Computational rankings support experimental prioritization. Experimental evidence remains the standard for biological validation.