MillionFull enables massive, full-length enzyme sequence-fitness data collection at low cost for machine learning-guided enzyme engineering

PRODUCTS USED

Genes
Read Full Article

ABSTRACT

Machine learning holds great promise for accelerating enzyme optimization, but its power is fundamentally constrained by the limited availability of sequence-fitness data. Here, we introduce MillionFull, a low-cost method that enables high-throughput full-length sequence- fitness mapping for enzymes of arbitrary length. Each run yields on the order of 10⁵-10⁷ data points, capturing sequence-function relationships at unprecedented scale. By overcoming the data bottleneck, MillionFull provides a foundation for dramatically advancing AI-driven enzyme engineering.

Read Full Article

PRODUCTS USED

Genes