AI Intelligence // signal over noise
← back to feed
Latent Space 7/10 signal

🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

researchindustry
Summary
Xaira Therapeutics has developed a new model, X-Cell, for predicting changes in gene expression, trained on a proprietary dataset called X-Atlas. This dataset was generated using CRISPR-based experiments to create causal, interventional data, which the company claims provides ~30x more information than existing observational datasets and overcomes previous model scaling limitations.
Context
Prior AI models for cell biology, often called "Virtual Cell" models, have largely been trained on observational datasets like the Chan Zuckerberg Institute's CELLxGENE, which contains data on 168 million cells. An influential example is scGPT, built by Bo Wang, who is now Xaira's Chief AI Scientist. While these models can describe correlations between cell types and states, they cannot predict the causal effects of interventions, such as a drug altering a specific gene. This creates an information bottleneck; models trained on this data see their test loss flatten after ~1.5B parameters, meaning more compute or larger models yield no improvement. Xaira's strategy is to break this wall by generating its own information-rich, causal data through large-scale biological experiments.
Details

Scaling Limitations with Observational Data:

  • Training on existing, smaller datasets revealed an information gap. A model's test loss would flatline after reaching 1.5B parameters, even as training loss continued to decrease.
  • A 3.1B parameter model fell off the scaling trend, indicating the model was limited by the information content of the data, not by parameters or compute.

Causal Data Generation (X-Atlas):

  • To overcome this, Xaira generated a new dataset called X-Atlas. The goal was to move from correlational data to causal data.
  • The data was created using CRISPR-based experiments that systematically perturb one gene at a time to observe the downstream effects on other genes (e.g., to determine if A causes B and C).
  • These experiments run millions of tests in parallel to build a map of causal relationships.
  • The author estimates the cost of data collection experiments and infrastructure was likely a "few tens of millions" of dollars, with compute and personnel costing a "few million" more, comparing the budget to a reinforcement learning rollout rather than typical pre-training.

The X-Cell Model:

  • X-Cell is the model trained on the X-Atlas dataset.
  • The architecture was changed from autoregression, used in prior models like scGPT, to a diffusion-based approach.
  • The model reportedly beats a linear baseline that had outperformed previous, more complex models.
  • It is designed to generalize to real-world lab experiments in human cells.
What's new
The primary innovation is the strategic shift from training models on existing, observational biological data to generating a large-scale, proprietary, interventional dataset (X-Atlas) specifically to uncover causal relationships. This 'data-first' approach, using high-throughput CRISPR experiments to build the training set, is what allows their model (X-Cell) to break through the performance scaling walls that limited previous models in the field.
Limitations
The source is a podcast summary, not a peer-reviewed paper, and lacks specific quantitative benchmark scores for X-Cell's performance. The cost estimates for data generation are explicitly stated as guesses by the podcast host. The source also notes that the current approach focuses on first-order effects (one gene changing at a time), while more complex second and third-order effects from multiple simultaneous gene changes remain a future challenge.
The take

Xaira's strategy is a powerful demonstration of the 'data as the moat' principle in a scientific domain. When public data hits an information ceiling, the frontier moves to those who can afford to generate proprietary, causal data. This is incredibly capital-intensive, suggesting the future of AI in fields like drug discovery will belong to vertically integrated companies that control both the wet lab (data generation) and the dry lab (model building). The comparison of their data budget to an 'RL rollout' is telling: this isn't just data collection, it's active, expensive experimentation to teach a model how the world works. This is the playbook for moving from correlational description to causal prediction.

Don't read this site daily. Get it in your inbox.

The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.