Cellular perturbation, modelled

Disease is a cell in the wrong state.

We build models that learn how cells move between states — and design the molecules that move them back.

Baseline state
Perturbed state
Fig. 01 — Cell-state manifold under perturbation
§ 01 — Thesis

A cell's identity is a point in a very high-dimensional space. Illness moves it. The right molecule moves it back.

Three things now exist at the same time: perturbation screens large enough to train on, architectures that transfer across cell types and contexts, and generative models that propose real chemical matter and real proteins. Biochemat is being built at that junction. The models are not the product — they are how we decide which molecules are worth making.

Read

Single-cell transcriptomics gives us a dense, genome-wide readout of what a cell is doing — millions of cells, one measurement at a time.

Simulate

Trained across genetic and chemical perturbation at scale, a model can propose how a cell will respond before the experiment is run.

Design

Invert the model and the question changes: not what will this molecule do, but what molecule produces the state we want.

§ 02 — Platform

Three models, one loop.

Each stage is a model in its own right. Together they close the loop from measurement to designed intervention.

01 Representation

A learned coordinate system for cell state

A cell is an unordered set of gene–expression pairs, not a sequence. So the model is permutation-equivariant attention over learned gene embeddings, trained by masked expression modelling. What we want from it is an embedding where distance carries biological meaning and where label transfer still works when labels are scarce — measured against strong baselines rather than instead of them.

masked expression modelling permutation-equivariant attention few-shot label transfer
scRNA-seq · 14 cells × 28 genes
expression zero / dropout held out predicted
Fig. 02 — Masked expression reconstruction
02 Simulation

The perturbatome, in silico

An atlas tells you how expression is distributed. A perturbation screen tells you how it is distributed under an intervention — and only the second grounds a claim about what a knockdown or a compound actually does. Our models map a control population to a perturbed one, so a prediction is a shift in a distribution, not a point estimate. The generalisation axis gets stated every time: unseen perturbation, unseen cell context, unseen combination, unseen dose. They are different problems.

Perturb-seq CRISPRi / CRISPRa chemical atlases distribution shift
response · control → perturbed
control perturbed predicted
Fig. 03 — Predicted distribution shift
03 Design

Molecules that produce a state

With a target state written in the model's own coordinates, design inverts: not what will this molecule do, but what produces this. We generate chemical matter conditioned on the response it should induce, with a retrosynthesis planner in the loop rather than a synthesisability score bolted on afterwards. For proteins, the capability that matters is not speed but epitope control — designing against a functionally relevant patch instead of whichever surface is easiest to raise a binder against. The map from signature to structure is one-to-many, so this is hypothesis generation. The bench stays the arbiter.

signature-conditioned generation epitope-conditioned binders flow matching retrosynthesis in the loop
noise σ 1.00
Fig. 04 — Epitope-conditioned binder, denoised
§ 03 — Methods

We pick the architecture the problem asks for.

Transformers do a great deal of the work, and they are not the answer to everything. Set-structured expression data, molecular graphs, three-dimensional structure and small-sample assay readouts each reward a different inductive bias.

Attention over gene sets

Expression is an unordered set, not a sequence. Attention with learned gene embeddings handles long-tail sparsity and variable panels.

Equivariant networks

For structure, SE(3)-equivariant message passing respects the symmetries of physical space instead of learning them from data.

Diffusion & flow matching

Generative trajectories for molecules and protein backbones, conditioned on a pocket, an epitope, or a target expression signature.

Latent-variable models

Where perturbation effects compose, structured latent spaces let interventions be added, transferred and disentangled from context.

Uncertainty, parameterized

Predictions are weighted by their own uncertainty, and experiments are chosen for how much of it they remove together.

Causal structure

Perturbation is intervention, not observation. Where the data supports causal identification, we use it rather than fitting correlations.

§ 04 — Direction

A platform is not the point. The point is the molecule at the end of it.

We are early, and precise about what that means: the models come first, the programmes follow. The design space is far larger than any lab can test, so the job of every model we build is to rank it — to decide which few experiments deserve the capacity. Everything points at one output: an intervention that moves a diseased cell measurably toward a healthy one, with the evidence to say so before it is made.

§ 05 — Contact

If you work on this, we should talk.

Scientists, engineers, collaborators with perturbation data, and people who want to build a drug company from the model outward.