Topic

ClinVar / TraitGym

All digests tagged ClinVar / TraitGym

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics) thumbnail

· 1:31:59

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

The video details the rapid evolution of Generative Genomics, focusing on how large language models (LLMs) trained on DNA sequences (Genome Language Models or GLMs) have advanced from merely reading DNA to actively designing functional biological sequences. Key models discussed include Hyena DNA, EVO, and the latest iteration, Omni. The core technical leap is Omni's ability to outperform specialized models across diverse tasks, such as predicting disease-causing mutations and understanding non-coding regulatory regions. This capability creates a dual mandate: advancing biological design while simultaneously developing advanced biosecurity tools to detect and counter engineered pathogens, framing the field as an AI arms race.

Key takeaways

  1. The Leap from Reading to Writing DNA

    Early models like Hyena DNA focused on reading DNA and predicting function using convolutions for long context (up to a million). Generative models like EVO marked the shift to generating sequences, culminating in the ability to generate functional genomes from scratch, a feat previously impossible for humans.

  2. Omni's Advancement via Alignment 18:30

    Omni represents a significant step beyond EVO by incorporating extensive mid-training and post-training (alignment). This process makes the pre-trained base model useful for specific scientific tasks, such as identifying causal variants, allowing it to outperform specialized models across a wide range of genomic tasks.

  3. The Biosecurity Arms Race 20:10

    The capability to design novel biological sequences necessitates a corresponding defensive capability. The defense must move beyond simple sequence matching and become function-aware, capable of detecting pathogens that look structurally different but maintain the same biological function.

  4. Mechanistic AI for Biology 29:10

    Mechanistic approaches involve probing the model's internal representations (embeddings and activations) to distill underlying biological patterns, such as GC content or transcription factor motifs. This allows researchers to understand the 'rules' the model has learned from the raw data.

Watch on YouTube Full article