🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Scaling Limitations with Observational Data:
- Training on existing, smaller datasets revealed an information gap. A model's test loss would flatline after reaching 1.5B parameters, even as training loss continued to decrease.
- A 3.1B parameter model fell off the scaling trend, indicating the model was limited by the information content of the data, not by parameters or compute.
Causal Data Generation (X-Atlas):
- To overcome this, Xaira generated a new dataset called X-Atlas. The goal was to move from correlational data to causal data.
- The data was created using CRISPR-based experiments that systematically perturb one gene at a time to observe the downstream effects on other genes (e.g., to determine if A causes B and C).
- These experiments run millions of tests in parallel to build a map of causal relationships.
- The author estimates the cost of data collection experiments and infrastructure was likely a "few tens of millions" of dollars, with compute and personnel costing a "few million" more, comparing the budget to a reinforcement learning rollout rather than typical pre-training.
The X-Cell Model:
- X-Cell is the model trained on the X-Atlas dataset.
- The architecture was changed from autoregression, used in prior models like scGPT, to a diffusion-based approach.
- The model reportedly beats a linear baseline that had outperformed previous, more complex models.
- It is designed to generalize to real-world lab experiments in human cells.
Xaira's strategy is a powerful demonstration of the 'data as the moat' principle in a scientific domain. When public data hits an information ceiling, the frontier moves to those who can afford to generate proprietary, causal data. This is incredibly capital-intensive, suggesting the future of AI in fields like drug discovery will belong to vertically integrated companies that control both the wet lab (data generation) and the dry lab (model building). The comparison of their data budget to an 'RL rollout' is telling: this isn't just data collection, it's active, expensive experimentation to teach a model how the world works. This is the playbook for moving from correlational description to causal prediction.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.