Latent Space

Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)

Published 2026-07-21 · Duration 1:29:47

Summary

Xaira Therapeutics introduced X-Cell, a novel 4.9-billion-parameter diffusion language model designed as a virtual cell foundation model of biology. The model is trained on X-Atlas/Pisces—a massive dataset spanning 25.6 million single cells across 16 biological contexts and generated via the Perturb-seq platform. The core breakthrough lies in shifting from descriptive (observational) data to causal (interventional) data, allowing the model to predict how a cell will respond to genetic perturbations it has never encountered. This capability is crucial for advancing drug discovery by moving beyond trial-and-error methods.

Download summary

Key takeaways

  1. Causality vs. Correlation in Biology 20:07

    Observational atlases (descriptive data) can describe biology, but they are fundamentally underpowered to learn causality. To predict the outcome of an intervention (e.g., knocking down a gene), causal data—generated through high-throughput perturbation screens—is required.

  2. X-Cell Architecture and Training 1:03:27

    X-Cell utilizes a diffusion language model approach, which treats gene expression prediction as an iterative 'editing' process rather than an autoregressive one. This architecture allows it to generate high-dimensional transcriptomic data by refining noisy representations until they minimize loss against the ground truth.

  3. Data Generation Scale and Engineering 1:16:47

    The model is powered by Perturb-seq, a technique combining high-throughput CRISPR perturbation with single-cell RNA sequencing. This process generates massive 2D datasets (perturbation on one axis, gene expression on the other) across millions of cells while minimizing batch effects.

  4. Generalization and Translational Potential 1:25:07

    X-Cell demonstrated impressive generalization by accurately predicting perturbation responses in active T-cells, even when the model was only trained on resting T-cell data. This suggests the potential to predict novel biology in unseen contexts.

Technical details

  • Virtual Cell Modeling 2307s

    A high-level concept for an AI model that predicts cell function or gene expression after specific interventions. The field aims to move beyond foundation models (which provide semantic representations) toward dynamic, predictive models.

  • Model Architecture Shift 3807s

    The development moved from autoregressive models (like scGPT), which assume an inherent order of genes, to diffusion language models. Diffusion models treat the prediction as an iterative refinement process, making them suitable for high-dimensional gene expression data where no natural linear order exists.

  • Data Inputs and Priors 4307s

    X-Cell incorporates a diverse set of biological priors to enhance context-specific predictions. These include literature embeddings, Protein-Protein Interaction (PPI) networks, and developmental map information (DevMap). The model can learn which prior is most important for specific cell types.

  • High-Throughput Data Generation 4607s

    The Perturb-seq technique uses CRISPR/Cas9 to systematically perturb (knock down) every gene in the human genome across millions of cells. This is coupled with single-cell RNA sequencing to read out the resulting impact on all 20,000 genes simultaneously.

Mentioned resources

  • X-Cell (Virtual Cell Model/AI Platform)
  • X-Atlas/Pisces (Genome-wide CRISPRi Perturb-seq Dataset)
  • scGPT (Single-cell Foundation Model (Predecessor))
  • PDB (Protein Structure Database)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.