ForetoData / research

ML-guided functional discovery and engineering.

Research program

Undercharacterized microbial systems provide the scientific setting. Comparative evidence, protein representations, multi-omic integration, and experimental prioritization provide the methods.

The program asks how machine learning can narrow biological search spaces without losing sight of the experiment that must eventually resolve uncertainty.

Flagship problem: oxygen-tolerant nitrogen fixation

Nitrogen fixation is an unusually clear test of this approach. Nitrogenase is oxygen-sensitive, while cyanobacteria produce oxygen through photosynthesis. Understanding how some organisms coordinate those processes requires functional discovery across genes, regulatory context, environmental condition, and phenotype.

Oxygen-tolerant nitrogen fixation is the flagship application. The same framework can support other problems where biological function is undercharacterized and experimental capacity is constrained.

A connected methodological program

Functional discovery

Begin with a biological function and the evidence required to distinguish a promising candidate from a convenient correlation.

Environmental and condition-specific biology

Expression and phenotype depend on context. Condition-specific transcriptomic and proteomic evidence helps define when a mechanism is relevant.

Sequence and protein representation

Comparative genomics and protein-language-model representations make it possible to reason across related genes and proteins even when annotation is incomplete.

Multi-omic integration

No single data layer settles the question. The task is to combine partially informative evidence while keeping its provenance visible.

Experimental prioritization

The output is a smaller, defensible set of candidates and a clear account of why each candidate is worth testing.

Iterative model-and-experiment cycles

Active learning treats each experimental result as the beginning of the next round, building a more efficient learning system over time.

Protein learning and engineering

For protein-property questions, the same logic moves from sequence representation through candidate selection to validation planning.

Figure 01 / iterative discovery system
Conceptual workflow. It distinguishes computational prioritization from experimental validation.

Computational rankings support experimental prioritization. Experimental evidence remains the standard for biological validation.