策略性临在:重构权力平衡
单细胞生物学正成为现代医学和药物研发的全新基石,其最引人瞩目的应用之一是细胞重编程(Cellular Rejuvenation: 通过表观遗传重塑使细胞恢复年轻状态)。2006年,山中伸弥(Shinya Yamanaka)发现了四种特定的转录因子——山中因子(Yamanaka Factors: Oct3/4、Sox2、Klf4、c-Myc 四种转录因子蛋白),通过过度表达这些因子,成功将老化的人类皮肤细胞重编程为类似胚胎干细胞的状态,并因此荣获2012年诺贝尔生理学或医学奖。这一发现为再生医学(Regenerative Medicine)开辟了全新篇章,使得再生组织或器官成为可能。在此基础上,局部重编程(Partial Reprogramming: 仅改变细胞的表观遗传年龄而不改变其细胞类型)技术被提出,旨在通过引入诸如 mRNA 药物来恢复细胞因衰老而丧失的年轻机能。在2026年,首款基于 OSK(Oct4, Sox2, Klf4)的重编程药物已经进入人体临床试验阶段,标志着这一技术从实验室走向临床历经了整整20年。
研究单细胞的另一大核心驱动力是构建一个统一、全局的虚拟细胞(Virtual Cell: 利用计算模型模拟单细胞的生理行为与反应)以及虚拟组织、虚拟器官直至“数字孪生人”。与计算机行业遵循的摩尔定律(Moore's Law: 芯片性能每隔约两年翻倍)相反,制药行业正面临所谓的反摩尔定律(Eroom's Law: 研发一款新药的成本呈指数级上升,而效率持续下降)。目前新药研发的失败率极高,最终获批率甚至不足 5%,整个研发管线动辄耗时10年以上且耗资数十亿美元。通过引入虚拟细胞,科学家有望在干湿实验交替中大幅缩短这一研发周期。然而,要实现这一愿景,创新不能仅局限于早期的蛋白质设计等单一环节,必须贯彻整个药物研发的完整生命周期。
在建立这种心理防线后,具体的语言博弈技巧如下,首先需要对单细胞进行多维度的分子测量。
Original English Source
My name is Akram. I'm a machine learning engineer at Altos Labs. Altos Labs is a biotech startup and the goal is to restore cell health and resilience through cellular rejuvenation to inverse disease and disabilities that can happen throughout the life. And the title of my talk is from tokens to cells. And this is my kind of view as someone without bio background to looking into the engineering challenges of foundation models for single-cell biology. And what I want to talk about first, what is single-cell? Why do we care about single-cell? How do we measure it? What are the problems with the data, getting the data? And then looking at current state of the art foundation models and then some takeaways in the end. Uh so, what is single-cell and why do we care about it? I would like to start with my favorite example, Yamanaka factor. In 2006, Shinya Yamanaka discovered four transcription factors. There are four specific type of a proteins that when they were overexpressed in a cell, they did it in a cell which was age old skin cell, it could reprogram the skin cell old skin cell to embryonic stem cell like state. So, basically from cell type skin cell old, it could reprogram it back to a young embryonic stem cell type. And this was a breakthrough for biology and then got him Nobel Prize later in 2012 because of the application and the new chapters possibilities for medicine. From regenerative medicine for like to be able to kind of regenerate tissues or organs. When we can like reprogram specific cell to any cell that we want to partially programming and aging application. For partial programming specifically is that we can also turn out that we can also only change the age of the cell. We don't have to change the type. So, by changing the age of the cell the goal is the hope is we can restore some of youthful function that we lose toward life as we age and then kind of like having this as a medicine maybe like an mRNA medicine for curing the disease that we encounter as we age. And that's a one example. And then I think this year 2026 we got the first type of reprogramming medicine. I think it's OSK and it's going to be tested in a human. So, it took kind of like 20 years. And the reason the other reason is that why we want to study a single cell is that we would like to model to have this unified holistic view of single cell to be able to model cell and from there hopefully we can model tissue and organ and then the entire human. And there are like projects like human cell atlas that actually started this effort this initiative and then they mapped every single cell inside human. And the goal ultimately is something like maybe for a ultimately some after some long time we can actually actually model like model human body or there are terminologies like virtual cell, virtual tissue, virtual human, digital twins. They all saying the same thing that the more we can model this living organisms. The better we are in understanding our body and how we can treat medicine, we can develop drugs. And the other problem is that so we you know, I think for this audience we know about Moore's law, the computer is getting like double every year. And then we have exact opposite on drug development. Basically, the number of drugs that developed each year is kind of declining, which is surprising with all the advances in technology, in AI. This is surprising to see. And then drug development is a field that failure is you know, very normal. Maybe the acceptance rate is kind of like you know, 5% or even less. And when we're looking at drug development pipeline from the early like research and development all the way to preclinical, clinical trial and then the final stage. The whole pipeline it could take up to 10 years easily and then it cost it can cost billions. And the goal here is that with the advances of AI and also like with this virtual cell, virtual organ, virtual human, the goal is that we can kind of like reduce this time, reduce this timeline. And also we're looking at the entire timeline. Let's say if we only look at like research and development, the models like you know, protein design and the stuff, we might like save like you know, few years here, but at the end maybe it's not going to help for the entire pipeline. So, it's important to have innovation and breakthrough across all pipeline.多模态单细胞测量与物理噪声挑战
为了构建高精度的虚拟细胞模型,首先需要对单细胞进行多维度的分子测量。在单细胞尺度上,主要包含以下几种测量模态:
- 基因组学(Genomics: 测量细胞内的 DNA 序列,在绝大多数同体细胞中基本保持一致)。
- 转录组学(Transcriptomics: 即单细胞 RNA 测序 (RNA-seq: 测量在特定时间点细胞内表达的 RNA 种类及其丰度),是目前单细胞大模型训练最主流的数据源)。由于聚合酶链式反应(PCR)等技术的普及,RNA-seq 数据极易扩展,目前的公开数据集规模已达数千万至数亿级别,甚至有项目在挑战十亿级细胞的制图。
- 蛋白质组学(Proteomics: 测量细胞内行使核心功能蛋白质的丰度)。尽管蛋白质是生理功能的主要执行者,但其测量技术通量极低(Low Throughput)且极其困难。
- 形态学(Morphology: 刻画细胞的三维形状和空间结构),通常依赖高分辨率显微成像技术,在结合空间转录组学以获取细胞在组织中的具体地理坐标(Spatial Dimension)时尤为关键。
然而,当前的单细胞测量数据面临巨大的工程与物理挑战。由于单细胞内部的分子转录过程往往以转录脉冲(Transcriptional Bursting: 基因表达并非连续,而是呈现阵发性波动的特征)的形式发生,且测量手段会对细胞造成不可逆的破坏,因此捕获的单细胞数据仅是细胞动态生命轨迹中的一个“静态瞬间”(Snapshot),而非连续的动态影像。此外,两颗基因型完全相同的细胞,其转录组读数也可能大相径庭,这使得单细胞数据表现出极高的噪声与异质性(Heterogeneity: 细胞由于微环境或随机波动导致的表达差异)。再加上不同实验室、不同仪器所产生的批次效应(Batch Effect: 非生物学因素带来的系统性数据偏差),使得数据对模型的泛化能力提出了极高要求。
在明确了单细胞数据的多模态及高噪声特征后,大模型技术在单细胞领域的具体探索路径如下。
Original English Source
Now that we know single cell is important, how can we measure single cell? So, looking at single, yeah, it's amazing. This is just one single cell. There is a lot going on inside that one little tiny organism. Uh there are different modalities and each modality is measure different things. Uh from genome, which is like, you know, DNA sequencing is almost, you know, same for all cells, to RNA sequence. RNA sequence or transcriptomic is like the profile of the genes that they are expressed at each time. And then that it's like a matrix and then basically cell it says for each cell uh what are the genes what are the genes how many genes are expressed and then in which quantity. And is like 20,000 K. It's it's 20,000 scale. And the other maybe modality is proteomics. Proteomics maybe is like what maybe we care about more like because the proteins are the one that they're doing most of the functions in our body. But the problem is that proteomics is a very hard to measure and you know, there are ways to measure it but it's a very low throughput. And morphology is about like the cell shape and structure. And there are like really good imaging technologies to kind of also measure to have like microscopic imaging from single cell. And it's very I've also useful for a special special dimension kind of like, you know, knowing the place of like, you know, how the single cell is located within tissue. Um so, from all these modalities maybe the one that has been used mostly for foundation model training, I would say is RNA-seq. And there is an issue that the technology is is it's easier to measure and the technology is like because of like PCR technologies and the stuff, it's easier to a scale. So we have like usually data set for single cell RNA in the scale of like tens of tens of uh millions of cells to even like you know, I've heard 500 million cells and then kind of like even like there's a project 1 billion cells. But usually what they talk about is RNA-seq. And then also this is also a problem that if you really want to understand biology and then starting from single cell, we really need to kind of have technological advances in other dimensions as well. And now having said that, let's say we got like you know, gathered 1 billion RNA-seq, 1 billion cells, 1 billion samples of RNA-seq data. Is this good for foundation models that we can you know, train on? And then the problem is uh it's it's very hard. The nature of the data is if you measure two identical cell at the they don't read the same. And then the problem is the problem is this is very very heterogeneous. I'll give you give an example here that usually cells they're living organism, they go through a lot of changes and cycles like from like you know, growing to uh kind of copying RNA to cell division. And usually some of these changes happens like in a burst, not like it's not something you know, continuous. And what we are measuring with current uh single cell sequencing is like a snapshot. We're taking a snapshots from from a whole movie. Uh now with that and then also there's their biological reason that the data is very noisy, it's very heterogeneous. Also, there is like technical reasons that usually if you measure in different labs, different machines, also the data would be the data wouldn't be the same.单细胞大模型:从Transformer到流匹配
在单细胞大模型的架构探索中,Transformer 架构(Transformer Architecture: 基于自注意力机制的深度学习模型)被率先引入。代表性工作如 scGPT 与 Informer 等。这些模型将单个“细胞”视为自然语言中的“句子”,而将“基因”视为“Token”。在预训练阶段,模型通常采用类似 BERT(Bidirectional Encoder Representations from Transformers: 双向编码器表征模型)的掩码机制(Masking),通过双向自注意力机制去预测被遮蔽的基因表达量(Gene Counts),以此学习基因与基因之间的复杂共表达关系。提取的隐空间表征被广泛应用于细胞类型分类、单细胞多模态整合以及扰动响应模拟(Perturbation Response Modeling)等下游任务。然而,在学术界(如 NeurIPS 上的多项基准测试)的实际评估中发现,此类大模型在进行编码时,会因为过度压缩而损失关键信息,导致其在许多下游任务上的表现甚至无法显著超越简单的线性模型,且其巨大的计算开销与产出不成正比。
为了克服自回归或自编码模型只倾向于预测表达均值而丢失分布特性的缺陷,研究人员提出了基于流匹配(Flow Matching: 一种生成式路径规划框架,通过常微分方程将简单噪声分布转换为复杂数据分布)的全新生成大模型——PrimeFlow。该模型直接从高斯噪声(Gaussian Noise)出发,通过拟合速度场来生成符合真实单细胞生理状态的概率分布。实验结果表明,PrimeFlow 在分布拟合指标(如 MMD 距离)上显著优于传统的自动编码器(Autoencoder)架构。未来的核心工程挑战在于,如何不仅在数据量(Data Volume)上进行堆叠,更要在数据采集质量上取得突破,以更真实地还原活体生物内部连续变化的动力学特征,从而实现真正意义上的单细胞基座模型泛化。
Original English Source
Now, let's see how what are look at some of the foundation models that we have and then see what they're doing with this uh, data. Um, and then also like maybe a small note here that given the nature that is very high dimensional is like multi-state. I think maybe, you know, some might argue that uh, a quantum computing would be there like kind of like natural fit for this application. And then, yeah, we don't know, it might be true. Uh, but you know, until we have quantum compute, for now we want to see what we can get with current state of the art AI. And now, uh, I want to start with like transformer-based models. And there have been really good like, you know, papers and models out there from the communities starting from like SCGPT and Informer. Uh, and what they want, they kind of like treat each uh, cell like a sentence and then genes like tokens. So, basically cell cell are built from genes. And what they do is like uh, uh, they um, like you know, BERT BERT style model like from language that kind of uh, masking some of the genes and they're trying having the model to predict that them predict that genes uh, that gene counts. So, looking at the data is like a matrix. Uh, and then for each cell we have like, you know, we have the count of uh, each genes. How many genes are, you know, how many genes are active, uh, per cell. And then it's trying to mask this to like, uh, attention, bidirectional attention across, and then kind of understand the relationship between, you know, between these genes. And, uh, this is how it's trained. And then like there are some, uh, uh, some some some downstream tasks that this model is used to. Uh, one is for, uh, so like, you know, predicting cell type, kind of predicting, uh, also, um, perturbation response modeling. Uh, but overall, when we look into these models, what they try to do, they try to kind of get the single cell data and then compress it in a kind of a latent vector and then do either decoding for like generative task or like classification for, uh, classic for other task. And then what happens is when you, uh, compress this data, we're losing a lot of information. It doesn't preserve the, it doesn't preserve the information. And then that's why when we're looking into this model, like, um, these models, we see that sometimes, sometimes like, uh, maybe, uh, simple linear models are on par or like sometimes even outperforming this, uh, these, uh, models that, you know, it's like, you know, complex models with a lot of compute is being used to train those. Um, then also we had like two papers last year at NeurIPS, we did like comprehensive benchmarking on uh, on these models. One is, uh, multimodal data for like imaging and RNA-seq, and the other one is for perturbation response modeling. So I have the links at the end of my talk, uh, go ahead and check them out, but like the final, the results, uh, the results, uh, what we found out was the Uh, that also is known in the community that even though these models are very expensive to train, uh, at the end they're not performing as well comparing to, you know, other domain language, uh, and and imaging. And this is our newer and then so transformer-based also we have like flow matching-based models that you you started to try to uh, predict the distribution of the data and uh, starting from like, you know, Gaussian noise and then trying to to match the distribution. And we have uh, this paper, uh, we have this model PrimeFlow uh, that is available on archive and it seems that these models they seem to do better. There is a better comparing to trans autoregressive-based models, um, because it's tries to kind of match the distribution predict the distribution rather than kind of, you know, understanding the understanding the data and compressing it into into uh, uh, into into a a latent vector. Um, and this is like some of the results that uh, we have. So, on the left we see like PrimeFlow it's the green the green dots are uh, ground truth label and then we see that PrimeFlow is kind of trying to match the distribution. Whereas other models like CPADR autoencoder-based is kind of like just mapping to the it's trying to predict the mean rather than understanding understanding the distribution. And then we have like MMD score at the bottom. Uh, so yeah, that was uh, my top three things to take away from this talk is that single cell is important and it's important to understand this. It has application for cellular rejuvenation. It's a path to like digital human and they're helping with drug development cycle. Then also we looked at at the different modalities to measure from single cell. RNA sequence data is the one that maybe it's we have it more available in a scale and also it's measuring something important in expression profile. But at the same at the same time that we would like to kind of get more data, work on the scale, but it's important that to work on quality of the data of the way that we are measuring this data to be a bit more realistic of the real organism than than just like you know scale the data the way it is. And then my final conclusion is that for now it seems flow matching models they're doing better for single cell data trying to match to the distribution. And also it to be able to scale these models that you know they can also do well on the data they haven't seen, they haven't trained on. We would need you know massive scaling scaling of the data I would say and then also the quality and the way that we measure data. And then with that yeah we have like three papers SCGENE scope, perturbation last year in NeurIPS and then PrimeFlow. You check them out and yeah, let me know if you have any questions. Thank you.📌 文中提及的人物和组织
公司/组织: Altos Labs