李飞飞:AI的下一个十年是空间智能,构建世界模型面临三大挑战 Best Partners TV 2025-11-12

李飞飞:AI的下一个十年是空间智能

11月10日,李飞飞亲自撰文,认为生成式AI(Generative AI: 能够生成文本、图像、音频等新内容的AI模型)的下一个战场是空间智能(Spatial Intelligence: 理解、推理和操作三维空间信息的能力)。

View/Hide Original English

On November 10th, Li Feifei personally authored an article, positing that the next frontier for generative AI is spatial intelligence.

在文章中,她首次系统性地解释了什么是空间智能,它为什么如此重要,以及如何构建能够解锁空间智能的世界模型(World Models: 能够理解、推理、生成并与复杂世界交互的AI模型)。

View/Hide Original English

In the article, she systematically explained for the first time what spatial intelligence is, why it is so important, and how to build world models capable of unlocking spatial intelligence.

李飞飞还指出了当前AI存在的致命缺陷:虽然AI掌握了海量的抽象知识,但是对于物理世界的常识和空间规律,它几乎一无所知。

View/Hide Original English

Li Feifei also pointed out a critical flaw in current AI: although AI possesses vast amounts of abstract knowledge, it is almost entirely ignorant of common sense and spatial rules in the physical world.

这种缺陷直接卡死了AI升级的大动脉。

View/Hide Original English

This deficiency directly obstructs the main artery for AI's advancement.

她敲响警钟,提出AI的下一个十年的真正突破不再是堆砌文字,而是要解锁空间智能,这才是连接感知、想象和行动的终极能力。

View/Hide Original English

She sounded the alarm, proposing that the true breakthrough for AI in the next decade will no longer be about accumulating text, but about unlocking spatial intelligence, as this is the ultimate capability that connects perception, imagination, and action.

文章发布之后,立即在社交平台引发了热议。

View/Hide Original English

After its publication, the article immediately sparked widespread discussion on social platforms.

图灵的愿景与当前AI的局限

1950年,当计算机还只是用来做自动运算和简单的逻辑判断时,艾伦·图灵(Alan Turing: 英国数学家、逻辑学家,计算机科学和人工智能的先驱)在《计算机器与智能》一文中,就提出了一个影响深远的问题:机器能思考吗?

View/Hide Original English

In 1950, when computers were only used for automatic calculations and simple logical judgments, Alan Turing, in his article "Computing Machinery and Intelligence," posed a profound question: Can machines think?

在那个时代,这个问题显得有些天马行空,但是图灵凭借着非凡的想象力洞察到,智能或许不必天生,而是可以通过技术构建出来。

View/Hide Original English

At that time, this question seemed somewhat fantastical, but Turing, with his extraordinary imagination, insightfully recognized that intelligence might not be innate but could be constructed through technology.

正是这个想法,催生了后来被称为人工智能的科学领域。

View/Hide Original English

It was this idea that gave rise to the scientific field later known as artificial intelligence.

75年过去,AI确实取得了翻天覆地的进步,尤其是最近几年,以大语言模型(Large Language Models - LLM: 基于海量文本数据训练,能理解和生成人类语言的AI模型)为代表的生成式AI,已经从实验室走进了我们的日常生活。

View/Hide Original English

Seventy-five years later, AI has indeed made revolutionary progress, especially in recent years, with generative AI, represented by large language models, moving from laboratories into our daily lives.

这些AI展现出了曾经被认为不可能的能力,但是如果我们冷静下来审视,就会发现当前的AI依然存在巨大的局限。

View/Hide Original English

These AIs have demonstrated capabilities once thought impossible, but if we calmly examine them, we will find that current AI still has significant limitations.

李飞飞在文章中提到了几个关键问题:首先,自主机器人的愿景还远未实现,我们在科幻电影里看到的机器人场景,至今还停留在实验室或者高度受限的场景中。

View/Hide Original English

Li Feifei mentioned several key issues in the article: firstly, the vision of autonomous robots is still far from being realized, and the robot scenarios we see in science fiction films are still confined to laboratories or highly restricted environments.

其次,AI在加速科学研究方面的潜力还没有完全释放出来,很多研究的关键步骤依然需要人类手动完成。

View/Hide Original English

Secondly, AI's potential in accelerating scientific research has not yet been fully unleashed, with many critical steps in research still requiring manual completion by humans.

最后,AI还无法真正理解并且赋能人类的创作者,无论是在分子化学、建筑设计还是电影建模等领域,AI都没能成为真正的帮手,因为它缺乏对空间的理解。

View/Hide Original English

Finally, AI is still unable to truly understand and empower human creators; whether in molecular chemistry, architectural design, or film modeling, AI has not become a true helper because it lacks an understanding of space.

空间智能:人类认知的基石

为什么会出现这些问题?李飞飞认为,核心原因在于当前的AI没有掌握空间智能,而这种能力恰恰是人类认知的基础。

View/Hide Original English

Why do these problems arise? Li Feifei believes the core reason is that current AI has not mastered spatial intelligence, a capability that is precisely the foundation of human cognition.

要搞清楚这一点,我们得先回到智能是如何进化的这个根本问题上。

View/Hide Original English

To understand this, we must first return to the fundamental question of how intelligence evolves.

提到智能,很多人会首先想到语言、逻辑、数学这些能力,但是实际上,空间智能才是更基础、更古老的智能形式。

View/Hide Original English

When intelligence is mentioned, many people first think of abilities like language, logic, and mathematics, but in reality, spatial intelligence is a more fundamental and ancient form of intelligence.

李飞飞在文章中把空间智能称为人类认知的脚手架,它就像建筑施工时的脚手架一样,支撑着我们对世界的感知、理解、推理和创造。

View/Hide Original English

In the article, Li Feifei refers to spatial intelligence as the scaffolding of human cognition; it supports our perception, understanding, reasoning, and creation of the world, much like scaffolding during building construction.

从进化的角度来看,动物们通过感官感知世界的简单行为,就已经开启了空间智能的进化之路。

View/Hide Original English

From an evolutionary perspective, the simple act of animals perceiving the world through their senses has already initiated the evolutionary path of spatial intelligence.

随着进化,动物的神经系统越来越复杂,逐渐形成了从感知到行动的核心循环。

View/Hide Original English

As evolution progressed, animal nervous systems became increasingly complex, gradually forming a core loop from perception to action.

许多科学家推测,正是这种循环驱动了智能的进化,最终让人类成为了既能感知、又能学习、还能思考和创造的物种。

View/Hide Original English

Many scientists speculate that this loop drove the evolution of intelligence, ultimately making humans a species capable of perceiving, learning, thinking, and creating.

而空间智能,就是这个循环的核心。

View/Hide Original English

And spatial intelligence is at the core of this loop.

对于我们每个人来说,空间智能其实是一种我们每天都在用,但是却不自觉的能力。

View/Hide Original English

For each of us, spatial intelligence is actually a capability we use every day without conscious awareness.

比如我们对空间布局的认知、对空间关系的实时判断、基于长期经验形成的空间直觉以及最基础的空间推理,它们背后是对复杂的空间感知、推理和预测,而这正是当前AI所最缺乏的。

View/Hide Original English

For example, our perception of spatial layout, real-time judgment of spatial relationships, spatial intuition formed from long-term experience, and the most basic spatial reasoning all rely on complex spatial perception, inference, and prediction, which is precisely what current AI most lacks.

更重要的是,空间智能还是人类文明进步的关键驱动力。

View/Hide Original English

More importantly, spatial intelligence is also a key driving force behind the progress of human civilization.

空间智能驱动人类文明进步的三个历史案例

李飞飞在文章中举了三个极具代表性的例子。

View/Hide Original English

Li Feifei cited three highly representative examples in the article.

第一个例子是古希腊学者埃拉托色尼(Eratosthenes: 古希腊数学家、地理学家、天文学家,首次计算出地球周长)测量地球周长。

View/Hide Original English

The first example is the ancient Greek scholar Eratosthenes measuring the Earth's circumference.

公元前3世纪,埃拉托色尼发现了一个有趣的现象:在夏至这天,埃及的赛伊尼地区,太阳会直射到深井的底部,这时的物体是没有影子的。

View/Hide Original English

In the 3rd century BCE, Eratosthenes observed an interesting phenomenon: on the summer solstice, in the region of Syene, Egypt, the sun would shine directly to the bottom of deep wells, and objects would cast no shadows.

而同一时间,在北方的亚历山大港,一根竖杆的影子会形成7度的夹角。

View/Hide Original English

At the same time, in Alexandria to the north, a vertical pole would cast a shadow forming a 7-degree angle.

凭借空间智能,埃拉托色尼把这两个看似无关的现象转化成了一个几何问题,他认为这7度的夹角,其实是因为地球是球形,两个地点在地球表面的纬度差异导致的。

View/Hide Original English

Utilizing spatial intelligence, Eratosthenes transformed these two seemingly unrelated phenomena into a geometric problem, reasoning that the 7-degree angle was due to the Earth being spherical and the latitudinal difference between the two locations on its surface.

于是,他通过测量亚历山大港到赛伊尼的距离,再乘以50,最终计算出地球的周长约为4万公里。

View/Hide Original English

Thus, by measuring the distance from Alexandria to Syene and multiplying it by 50, he ultimately calculated the Earth's circumference to be approximately 40,000 kilometers.

这个结果和现代科技测量的实际值惊人地接近。

View/Hide Original English

This result is astonishingly close to the actual value measured by modern technology.

第二个例子是詹姆斯·哈格里夫斯(James Hargreaves: 英国发明家,珍妮纺纱机的发明者)发明珍妮纺纱机(Spinning Jenny: 18世纪中期英国发明的多锭纺纱机,显著提高了纺纱效率)。

View/Hide Original English

The second example is James Hargreaves inventing the Spinning Jenny.

18世纪中期,英国的纺织业还是手工操作,一个工人用传统纺纱机一次只能纺一根线,效率很低。

View/Hide Original English

In the mid-18th century, the British textile industry was still manual; a worker using a traditional spinning wheel could only spin one thread at a time, which was very inefficient.

哈格里夫斯在观察妻子纺纱的时候,突然有了一个空间层面的灵感:如果把多个纺锤并排排列在同一个框架里,是不是就能让一个工人同时纺多根线呢?

View/Hide Original English

While observing his wife spinning, Hargreaves suddenly had a spatial insight: what if multiple spindles were arranged side by side in the same frame, allowing one worker to spin multiple threads simultaneously?

基于这个想法,他发明了珍妮纺纱机,把纺纱效率提高了8倍。

View/Hide Original English

Based on this idea, he invented the Spinning Jenny, increasing spinning efficiency eightfold.

这个发明也直接推动了工业革命的爆发,让人类从手工生产进入了机器生产的时代。

View/Hide Original English

This invention also directly spurred the Industrial Revolution, transitioning humanity from manual production to the era of machine production.

第三个例子是詹姆斯·沃森(James Watson: 美国分子生物学家,DNA双螺旋结构的共同发现者)和弗朗西斯·克里克(Francis Crick: 英国分子生物学家、物理学家,DNA双螺旋结构的共同发现者)发现DNA双螺旋结构

View/Hide Original English

The third example is James Watson and Francis Crick discovering the DNA double helix structure.

20世纪50年代,科学家们已经知道DNA是遗传物质,但是不知道它的结构是什么样的。

View/Hide Original English

In the 1950s, scientists already knew that DNA was genetic material, but its structure remained unknown.

沃森和克里克没有依赖复杂的仪器,而是通过构建3D分子模型来寻找答案。

View/Hide Original English

Watson and Crick did not rely on complex instruments but sought answers by building 3D molecular models.

他们用金属板代表DNA的碱基,用金属丝代表连接它们的化学键,不断调整这些组件的空间排列,直到找到符合所有实验数据的结构。

View/Hide Original English

They used metal plates to represent DNA bases and metal wires to represent the chemical bonds connecting them, continuously adjusting the spatial arrangement of these components until they found a structure that matched all experimental data.

最终,他们发现DNA是由两条反向平行的链组成的双螺旋结构。

View/Hide Original English

Ultimately, they discovered that DNA is a double helix structure composed of two antiparallel strands.

这个发现揭开了生命遗传的奥秘,为现代分子生物学奠定了基础。

View/Hide Original English

This discovery unveiled the mysteries of genetic inheritance and laid the foundation for modern molecular biology.

这三个例子跨越了几千年的历史,但是都指向了同一个结论:空间智能是人类理解世界、改造世界的核心能力。

View/Hide Original English

These three examples span thousands of years of history, but all point to the same conclusion: spatial intelligence is humanity's core ability to understand and transform the world.

无论是探索自然规律、发明新工具,还是推动科学革命,都离不开对空间关系的感知、推理和创造。

View/Hide Original English

Whether exploring natural laws, inventing new tools, or driving scientific revolutions, all depend on the perception, reasoning, and creation of spatial relationships.

而当前的AI之所以无法在这些领域发挥作用,正是因为它缺乏这种能力。

View/Hide Original English

The reason current AI cannot function in these areas is precisely because it lacks this capability.

构建世界模型:AI的未来解决方案

那么,要想让AI拥有空间智能,我们需要做什么?

View/Hide Original English

So, what do we need to do to enable AI to possess spatial intelligence?

李飞飞和她的团队提出了一个关键的解决方案:构建世界模型。

View/Hide Original English

Li Feifei and her team proposed a key solution: building world models.

这种模型的核心能力是能够理解、推理、生成并且与语义、物理、几何和动态上都极为复杂的世界进行交互,这远远超出了当前大语言模型的能力范围。

View/Hide Original English

The core capability of such models is to understand, reason, generate, and interact with a world that is extremely complex semantically, physically, geometrically, and dynamically, far exceeding the capabilities of current large language models.

那么,一个真正的世界模型需要具备哪些核心能力?李飞飞在文章中明确提出了三点。

View/Hide Original English

So, what core capabilities must a true world model possess? Li Feifei clearly outlined three points in the article.

第一个核心能力是生成式(Generative):它能根据语义或者感知指令,生成无穷无尽、多样化的模拟世界,而且这些世界必须满足感知一致性、几何一致性和物理一致性。

View/Hide Original English

The first core capability is generative: it must be able to generate endless, diverse simulated worlds based on semantic or perceptual instructions, and these worlds must satisfy perceptual, geometric, and physical consistency.

在这个方面,目前学术界正在探索的一个关键问题是:世界模型应该如何表示这些空间信息?

View/Hide Original English

In this regard, a key question currently being explored in academia is: how should world models represent this spatial information?

是用隐式表示,通过神经网络的参数间接存储空间信息,还是用显式表示,直接存储物体的3D坐标、形状、物理属性呢?

View/Hide Original English

Should it use implicit representation, indirectly storing spatial information through neural network parameters, or explicit representation, directly storing objects' 3D coordinates, shapes, and physical properties?

李飞飞对此的观点是,两者都需要。

View/Hide Original English

Li Feifei's view on this is that both are needed.

第二个核心能力是多模态(Multimodal):我们人类在理解世界的时候,会同时使用视觉、嗅觉、触觉、听觉等多种感官。

View/Hide Original English

The second core capability is multimodal: when humans understand the world, we simultaneously use multiple senses such as sight, smell, touch, and hearing.

这种多模态融合的能力是我们理解世界的关键。

View/Hide Original English

This ability to integrate multiple modalities is crucial for our understanding of the world.

同样,世界模型也需要具备这种能力。

View/Hide Original English

Similarly, world models also need to possess this capability.

当前的多模态模型虽然也能处理多种输入,但是它的核心还是语言模型,图像、视频等模态只是辅助输入,并没有真正与语言模态深度融合。

View/Hide Original English

Although current multimodal models can process various inputs, their core remains a language model, with modalities like images and videos serving only as auxiliary inputs, not truly deeply integrated with the language modality.

而世界模型的设计理念是原生多模态,也就是说,各种模态在模型中是平等的,它们共同服务于理解和构建世界这个核心目标。

View/Hide Original English

The design philosophy of world models, however, is natively multimodal, meaning that various modalities are equal within the model, and they collectively serve the core goal of understanding and building the world.

第三个核心能力是交互性(Interactive):如果我们向模型输入动作或目标,它应该能输出世界的下一个状态,而且这个状态必须符合世界的历史状态、物理规律和语义逻辑。

View/Hide Original English

The third core capability is interactivity: if we input an action or goal into the model, it should be able to output the next state of the world, and this state must conform to the world's historical state, physical laws, and semantic logic.

这是世界模型最关键的能力之一,因为它直接对应人类的行动能力。

View/Hide Original English

This is one of the most critical capabilities of world models, as it directly corresponds to human agency.

而更高级的交互性能力还包括目标驱动的动作规划,也就是说,世界模型不仅要能预测目标的状态,还能规划出需要做哪些动作才能达到这个目标。

View/Hide Original English

More advanced interactive capabilities also include goal-driven action planning, meaning that a world model must not only predict the state of a goal but also plan what actions are needed to achieve that goal.

这也是未来机器人自主行动的核心基础。

View/Hide Original English

This is also the core foundation for future autonomous robotic actions.

如果机器人能够通过世界模型规划动作,那么它就能在复杂环境中自主完成任务,而不需要人类重新编程。

View/Hide Original English

If robots can plan actions through world models, they will be able to autonomously complete tasks in complex environments without requiring human reprogramming.

李飞飞认为,构建具备这三大能力的世界模型是AI下一个十年的决定性挑战,因为这个挑战的难度远远超过了构建大语言模型。

View/Hide Original English

Li Feifei believes that building world models with these three capabilities is the decisive challenge for AI in the next decade, because the difficulty of this challenge far exceeds that of building large language models.

因为语言是一维的、序列式的信号,而世界是三维的、动态的、多模态的系统。

View/Hide Original English

This is because language is a one-dimensional, sequential signal, whereas the world is a three-dimensional, dynamic, and multimodal system.

要模拟这样的系统,需要解决一系列技术难题。

View/Hide Original English

Simulating such a system requires solving a series of technical challenges.

李飞飞创立的World Labs就是专注于世界模型研究的机构。

View/Hide Original English

World Labs, founded by Li Feifei, is an institution dedicated to world model research.

构建世界模型面临的三大核心挑战

在文章中,李飞飞分享了当前世界模型研究面临的三大核心挑战,以及团队的一些探索成果。

View/Hide Original English

In the article, she shared the three core challenges currently facing world model research, along with some of her team's exploration results.

挑战一:设计通用的训练任务函数

View/Hide Original English

Challenge One: Designing a Universal Training Task Function

我们知道,大语言模型之所以能取得成功,一个关键原因是它有一个简单、优雅且通用的训练任务:预测下一个token(Token: 在自然语言处理中,文本被分割成的最小有意义单元)。

View/Hide Original English

We know that a key reason for the success of large language models is their simple, elegant, and universal training task: predicting the next token.

而世界模型的训练,目前还没有这样一个通用的任务函数。

View/Hide Original English

However, for training world models, such a universal task function does not yet exist.

为什么设计这个任务函数这么难?因为世界模型需要处理的是空间+时间+多模态的复杂系统,而不是一维文本。

View/Hide Original English

Why is designing this task function so difficult? Because world models need to process complex systems involving space + time + multiple modalities, rather than one-dimensional text.

这涉及到的连续性的维度更多,也更复杂。

View/Hide Original English

This involves more continuous dimensions, and they are more complex.

李飞飞认为,世界模型的训练任务函数必须反映几何和物理规律,因为这是世界的本质属性。

View/Hide Original English

Li Feifei believes that the training task function for world models must reflect geometric and physical laws, as these are the inherent properties of the world.

一个可能的任务函数也许是预测下一个世界的状态预测,但是这个任务函数的具体设计还需要进一步的探索。

View/Hide Original English

A possible task function might be predicting the next state of the world, but the specific design of this task function requires further exploration.

如何定义世界状态?如何衡量预测的准确性?这些都需要研究人员不断尝试和优化。

View/Hide Original English

How should the world state be defined? How should prediction accuracy be measured? These questions require continuous experimentation and optimization by researchers.

挑战二:如何获取大规模、高质量的训练数据

View/Hide Original English

Challenge Two: How to Obtain Large-Scale, High-Quality Training Data

大语言模型的训练数据主要是互联网上的文本数据,这些数据数量庞大、获取成本低。

View/Hide Original English

The training data for large language models primarily consists of text data from the internet, which is vast in quantity and low in acquisition cost.

而世界模型的训练数据需要包含空间、物理、多模态的信息,获取难度要大得多。

View/Hide Original English

However, training data for world models needs to include spatial, physical, and multimodal information, making it much more difficult to acquire.

当前的训练数据来源主要有三个:第一个是互联网上的图像和视频数据,比如YouTube、Instagram上有大量的视频,这些视频包含了丰富的空间和动态信息。

View/Hide Original English

Current sources of training data primarily include three categories: the first is image and video data from the internet, such as the vast number of videos on YouTube and Instagram, which contain rich spatial and dynamic information.

但问题是,这些数据大多是二维的,而且缺乏深度、触觉、物理属性等关键信息。

View/Hide Original English

However, the problem is that most of this data is two-dimensional and lacks critical information such as depth, tactile feedback, and physical properties.

所以,研究的关键是如何从二维图像/视频中提取三维空间信息。

View/Hide Original English

Therefore, the key to research is how to extract three-dimensional spatial information from two-dimensional images/videos.

第二个是合成数据,也就是通过计算机模拟生成的、带有精确标注的空间数据。

View/Hide Original English

The second source is synthetic data, which refers to spatially annotated data generated through computer simulations.

比如,用3D引擎生成大量的虚拟场景,并且标注出每个物体的3D坐标、物理属性、动态规律。

View/Hide Original English

For example, generating a large number of virtual scenes using a 3D engine and labeling each object's 3D coordinates, physical properties, and dynamic rules.

这种数据的优点是标注精确、可控性强,可以弥补真实数据的不足。

View/Hide Original English

The advantage of this type of data is its precise annotation and strong controllability, which can compensate for the shortcomings of real-world data.

但是合成数据也有缺点,它可能与真实世界存在着差距,导致模型在真实场景中表现不佳。

View/Hide Original English

However, synthetic data also has drawbacks; it may have discrepancies with the real world, leading to suboptimal model performance in real-world scenarios.

第三个是多模态的补充数据,比如通过深度相机获取的深度图、通过触觉传感器获取的触觉数据、通过运动捕捉设备获取的动作数据。

View/Hide Original English

The third source is complementary multimodal data, such as depth maps obtained from depth cameras, tactile data from tactile sensors, and motion data from motion capture devices.

这些数据能为世界模型提供更丰富的空间信息,但是获取的成本很高。

View/Hide Original English

This data can provide richer spatial information for world models, but its acquisition cost is high.

李飞飞认为,未来世界模型的训练数据必然是真实数据+合成数据+多模态数据的融合。

View/Hide Original English

Li Feifei believes that future training data for world models will inevitably be a fusion of real data, synthetic data, and multimodal data.

而研究的关键是如何构建能高效利用这些数据的模型架构,如何让模型从真实数据中学习世界的多样性,从合成数据中学习精确的空间和物理规律,以及从多模态数据中学习跨模态的关联关系。

View/Hide Original English

The key to research is how to build model architectures that can efficiently utilize this data, how to enable models to learn the world's diversity from real data, precise spatial and physical laws from synthetic data, and cross-modal relationships from multimodal data.

挑战三:需要研发新的模型架构和表示学习方法

View/Hide Original English

Challenge Three: The Need for New Model Architectures and Representation Learning Methods

当前的多模态和视频生成模型大多是基于Transformer架构(Transformer Architecture: 一种基于自注意力机制的神经网络架构,广泛应用于自然语言处理和计算机视觉),并且把图像、视频转化为一维或二维的token序列来处理。

View/Hide Original English

Current multimodal and video generation models are mostly based on the Transformer architecture and process images and videos by converting them into one-dimensional or two-dimensional token sequences.

这种方法有一个致命的缺点:它破坏了空间的三维结构和时间的动态连续性,导致模型无法真正理解空间关系。

View/Hide Original English

This method has a fatal flaw: it destroys the three-dimensional structure of space and the dynamic continuity of time, preventing the model from truly understanding spatial relationships.

因此,世界模型需要研发超越当前范式的新型架构。

View/Hide Original English

Therefore, world models require the development of new architectures that go beyond current paradigms.

李飞飞提到了两种可能的方向:第一种是3D/4D感知的tokenization方法,也就是说,不再把图像、视频转化为2D token,而是转化为3D空间token或者4D时空token。

View/Hide Original English

Li Feifei mentioned two possible directions: the first is 3D/4D-aware tokenization methods, meaning that images and videos would no longer be converted into 2D tokens but into 3D spatial tokens or 4D spatiotemporal tokens.

这样,模型就能直接处理空间的三维结构和时间的动态变化,而不是间接通过2D像素来推测。

View/Hide Original English

This way, the model could directly process the three-dimensional structure of space and dynamic changes over time, rather than inferring them indirectly through 2D pixels.

第二种是空间记忆机制,也就是说,模型需要有专门的记忆模块,用来存储世界的空间信息,比如房间里有哪些物体,这些物体的位置、形状、物理属性是什么。

View/Hide Original English

The second is a spatial memory mechanism, meaning that the model needs to have dedicated memory modules to store spatial information about the world, such as what objects are in a room, and what their positions, shapes, and physical properties are.

这样,当模型处理动态场景时,能通过调用空间记忆,保持场景的一致性。

View/Hide Original English

This way, when the model processes dynamic scenes, it can maintain scene consistency by recalling spatial memories.

World Labs的团队已经在这方面做了一些探索,比如他们最近研发的实时生成帧模型(Real-Time Generative Frame-based Model - RTFM)。

View/Hide Original English

The World Labs team has already made some explorations in this area, such as their recently developed Real-Time Generative Frame-based Model (RTFM).

这个模型的核心创新是用基于空间的帧作为空间记忆,也就是说,模型会把生成的世界场景存储为一系列带有空间坐标的帧。

View/Hide Original English

The core innovation of this model is the use of spatially-based frames as spatial memory; that is, the model stores generated world scenes as a series of frames with spatial coordinates.

当需要生成下一个状态时,模型会调用这些帧,确保空间的一致性。

View/Hide Original English

When the next state needs to be generated, the model calls upon these frames to ensure spatial consistency.

目前,RTFM已经能实现高效的实时生成,并且保持生成世界的空间一致性,这是世界模型研究的一个重要进展。

View/Hide Original English

Currently, RTFM can achieve efficient real-time generation while maintaining spatial consistency in the generated world, which is a significant advancement in world model research.

世界模型在Marble平台上的初步应用

除了技术挑战,World Labs还展示了世界模型的第一个应用产品——Marble平台(Marble Platform: World Labs开发的一款面向创作者的世界模型工具,能生成并维护一致的3D环境)。

View/Hide Original English

In addition to technical challenges, World Labs also showcased the first application product of world models: the Marble platform.

这是一个面向创作者的世界模型工具,能通过多模态的输入生成并且维护一致的3D环境。

View/Hide Original English

This is a world model tool for creators that can generate and maintain consistent 3D environments through multimodal input.

目前,Marble已经向部分用户开放测试。

View/Hide Original English

Currently, Marble has been opened for testing to a select group of users.

李飞飞博士表示,团队正在努力优化,争取尽快向公众开放。

View/Hide Original English

Dr. Li Feifei stated that the team is working diligently to optimize it and aims to open it to the public as soon as possible.

空间智能的未来应用前景

聊了这么多技术细节,可能有朋友会问,空间智能和世界模型最终能给我们的生活带来什么改变呢?

View/Hide Original English

After discussing so many technical details, some friends might ask, what changes can spatial intelligence and world models ultimately bring to our lives?

李飞飞在文章中从短期、中期、长期三个时间维度,描绘了空间智能的应用前景。

View/Hide Original English

In the article, Li Feifei outlined the application prospects of spatial intelligence across three time dimensions: short-term, medium-term, and long-term.

空间智能的第一个重要应用领域是创造力领域。

View/Hide Original English

The first important application area for spatial intelligence is creativity.

李飞飞认为,空间智能将彻底改变我们创造和体验叙事的方式,从传统二维的叙事方式到轻松生成可探索的3D叙事世界。

View/Hide Original English

Li Feifei believes that spatial intelligence will fundamentally change how we create and experience narratives, moving from traditional two-dimensional storytelling to effortlessly generating explorable 3D narrative worlds.

设计流程也将迎来效率革命,有了空间智能工具,设计师可以通过文本+草图的方式快速生成3D模型,进行实时调整。

View/Hide Original English

Design processes will also undergo an efficiency revolution; with spatial intelligence tools, designers can quickly generate 3D models using text and sketches, allowing for real-time adjustments.

沉浸式体验也会更加普及,打破空间的限制,让人与人之间的连接更加紧密。

View/Hide Original English

Immersive experiences will also become more widespread, breaking spatial limitations and fostering closer connections between people.

空间智能的第二个重要应用领域是机器人领域。

View/Hide Original English

The second important application area for spatial intelligence is robotics.

李飞飞认为,机器人要成为人类的协作伙伴,必须具备空间智能,因为机器人需要在物理世界中移动、操作物体、与人类互动,这些都离不开对空间的理解。

View/Hide Original English

Li Feifei believes that for robots to become collaborative partners with humans, they must possess spatial intelligence, as robots need to move, manipulate objects, and interact with humans in the physical world, all of which depend on an understanding of space.

当前的机器人技术最大的瓶颈就是缺乏空间智能,而世界模型的出现会从三个方面推动机器人技术的突破。

View/Hide Original English

The biggest bottleneck in current robotics technology is the lack of spatial intelligence, and the emergence of world models will drive breakthroughs in robotics technology in three aspects.

第一个突破是解决机器人训练数据稀缺的问题,世界模型可以模拟无数个虚拟训练场景,让机器人在虚拟环境中快速学习。

View/Hide Original English

The first breakthrough is solving the problem of scarce robot training data; world models can simulate countless virtual training scenarios, allowing robots to learn quickly in virtual environments.

第二个突破是让机器人成为懂人类的协作伙伴,未来的机器人需要具备空间智能+语义理解的能力,才能理解人类的需求和意图。

View/Hide Original English

The second breakthrough is enabling robots to become collaborative partners who understand humans; future robots will need to possess both spatial intelligence and semantic understanding capabilities to comprehend human needs and intentions.

第三个突破是拓展机器人的形态和应用场景,当前的机器人大多还是人形机器人或工业机械臂,形态比较单一。

View/Hide Original English

The third breakthrough is expanding the forms and application scenarios of robots; current robots are mostly humanoid robots or industrial robotic arms, with relatively singular forms.

而未来会根据应用场景的需求,呈现出多样化的形态。

View/Hide Original English

In the future, they will take on diverse forms according to the needs of different application scenarios.

除了创造力和机器人领域,空间智能在科学研究、医疗健康、教育等长期领域也将产生深远的影响。

View/Hide Original English

Beyond creativity and robotics, spatial intelligence will also have a profound impact on long-term fields such as scientific research, healthcare, and education.

在科学研究领域,空间智能将成为加速发现的工具,比如科学家可以用世界模型模拟全球气候系统的三维动态变化,或者模拟原子的空间排列与材料性能的关系,甚至是模拟星系的三维结构和演化过程。

View/Hide Original English

In the field of scientific research, spatial intelligence will become a tool for accelerating discovery; for example, scientists can use world models to simulate the three-dimensional dynamic changes of global climate systems, or the relationship between atomic spatial arrangements and material properties, or even the three-dimensional structure and evolution of galaxies.

这些模拟的核心都是空间关系和动态规律,而这正是世界模型的优势所在。

View/Hide Original English

The core of these simulations lies in spatial relationships and dynamic laws, which is precisely where world models excel.

通过空间智能,科学家可以突破实验条件的限制,探索那些无法在实验室中复现的现象。

View/Hide Original English

Through spatial intelligence, scientists can overcome experimental limitations and explore phenomena that cannot be replicated in laboratories.

在医疗健康领域,空间智能将成为提升诊疗水平的助手。

View/Hide Original English

In the healthcare sector, spatial intelligence will serve as an assistant to enhance diagnostic and treatment levels.

李飞飞在斯坦福大学的团队已经在这个领域做了多年研究,比如加速药物的研发、提升诊断的准确性以及优化患者的照顾。

View/Hide Original English

Li Feifei's team at Stanford University has been conducting research in this area for years, such as accelerating drug development, improving diagnostic accuracy, and optimizing patient care.

在教育领域,空间智能将成为革新学习方式的工具。

View/Hide Original English

In the field of education, spatial intelligence will become a tool for revolutionizing learning methods.

空间智能可以让学习变得更加沉浸式、交互式,让学习不再是死记硬背,而是亲身体验、主动探索,大大提高学习的效率和兴趣。

View/Hide Original English

Spatial intelligence can make learning more immersive and interactive, transforming it from rote memorization into hands-on experience and active exploration, thereby greatly enhancing learning efficiency and interest.

结语

在文章的最后,李飞飞提到,她投身AI领域25年,始终被图灵的愿景所激励着。

View/Hide Original English

In the conclusion of the article, Li Feifei mentioned that she has been dedicated to the field of AI for 25 years, consistently inspired by Turing's vision.

而空间智能的研究正是基于这个愿景,它不是要让AI取代人类的创造力,而是要让AI成为人类的伙伴,帮助我们突破自身的局限,实现以前无法实现的目标。

View/Hide Original English

The research into spatial intelligence is based on this vision; it is not about AI replacing human creativity, but about AI becoming a partner to humanity, helping us overcome our limitations and achieve goals previously unattainable.

最后,我想引用李飞飞的一句话作为结尾:她说,从自然界赋予远古动物空间智能的第一缕曙光开始,已经过去了将近5亿年。

View/Hide Original English

Finally, I would like to conclude with a quote from Li Feifei: she said that nearly 500 million years have passed since nature bestowed the first glimmer of spatial intelligence upon ancient animals.

我们有幸成为这样一代的技术人员,不仅可能很快就能赋予机器同样的能力,并且有幸利用这些能力为世界各地的人们带来福祉。

View/Hide Original English

We are fortunate to be a generation of technologists who may soon be able to endow machines with the same capabilities, and have the privilege of using these capabilities to bring well-being to people around the world.

📌 文中提及的人物和组织

人物: 李飞飞, Alan Turing

公司/组织: 斯坦福大学

产品/模型: DNA双螺旋结构, Marble平台

媒体/书籍: 《计算机器与智能》