AI能否发明广义相对论?智能的界限
Adam Brown: 或许,这些系统,这些大语言模型(LLMs)最终能够做到的最后一件事,就是在我们上个世纪之交所理解的物理定律的基础上,发明广义相对论。从这一点出发,我认为这可能就是终极步骤了。一旦它能做到这一点,那么就对人类而言,就没有太多其他事情可做了。
Original English
Adam Brown: maybe the very last thing that these systems will be able to do these llms will be able to do is given the laws of physics as we understood them at the turn of the last century invent general relativity
Adam Brown: 这相当非凡。我是说,尤其对于一个物理学背景的人来说,物理学领域的进展是相当缓慢的。但来到人工智能领域,却看到日复一日、周复一周、年复一年的惊人快速进展。
Original English
Adam Brown: it's pretty extraordinary I mean particularly coming from a physics background in which progress is pretty slow uh to come to the AI field and see progress being so extraordinarily rapid day by day uh week by week year by year
Adam Brown: 嗯,看着这一切,确实看起来这些大语言模型,以及这些人工智能系统,在某种意义上它们只是插值器。但它们进行插值的抽象层级却不断在提高,我们也在不断地沿着这个抽象链条向上攀升。然后,很可能从一个足够高的视角来看,从牛顿物理学中衍生出的生成能力,不过是在某种足够宏大的抽象层级上的插值,这或许能告诉我们关于智能本质、人类智能,以及这些大型语言模型的某些东西。
Original English
Adam Brown: um looking at it it certainly looks like these llms uh and these AI systems in some sense they're just interpolators but the level of abstraction at which they're interpolating keeps going up and up and up uh and we keep sort of writing up that chain of abstractions and then presumably from a sufficiently elevated point of view the invention of generativity uh from Newtonian physics is just interpolation at some sufficiently grandiose level of abstraction that perhaps tells us something about the nature of intelligence human intelligence as well as uh as well as about these large language models
Adam Brown: 如果你问我需要多少年才能做到这一点,那确实不完全清楚。但从某种意义上说,广义相对论是人类有史以来最伟大的飞跃。一旦我们能做到这一点,或许在 10 年内,我们就会完全涵盖人类智能。
Original English
Adam Brown: if you ask me how many years until we can do that uh that is not totally clear but um in some sense General general relativity was the greatest leap that Humanity ever made and uh once we can do that perhaps in 10 years uh then then we will have fully encompassed human intelligence
Adam Brown: 它会拥有与爱因斯坦相同的特质吗?显然,人类智能与这些大型语言模型之间存在许多类比上的不足。但我认为,在正确的抽象层面上,它们可能是相同的。
Original English
Adam Brown: will it have the same will it be of the same character as what Einstein did clearly there's some there are many disanalogies between human intelligence in these large language models but I think at the right level of distraction it it may be the same
Adam Brown: 你认为 AI 数学家和物理学家是否会比人类具有优势,仅仅因为它们默认就能以人类不擅长的方式思考奇怪的维度和流形?
Original English
Adam Brown: do you think AI mathematicians physicist will have advantages over humans just because they can by default think in terms of weird dimensions and manifolds in a way that doesn't natively come to humans ah
Adam Brown: 嗯,你知道,我认为我们可能需要回顾一下,人类在多大程度上是或不是天生地在高维空间中思考。这显然不是我们的自然空间。曾有一种技术被发明出来用来思考这些事情,那就是符号(notation)、张量符号(tensor notation)等等,这些东西让你能够……即使是通过爱因斯坦 100 年前的写作方式,也能很自然地在维度之间转换。然后你就更多地在思考如何操作这些数学对象,而不是在高维空间中思考。
Original English
Adam Brown: um you know I think maybe we need to back up to in what sense the humans do or don't think natively in high Dimensions obviously it's not our natural space there was a technology that was invented to think about these things which was you know notation tensor notation VAR other things that allows you to much using just even writing as as Einstein did 100 years ago allows you to of naturally move between Dimensions uh and then you're thinking more about manipulating these mathematical objects than you are about thinking in higher Dimensions
Adam Brown: 我认为,大型语言模型在思考高维空间方面,并不比人类更自然。你可以说,大型语言模型拥有数十亿个参数,这就好比一个十亿维度的空间。但你也可以对人类大脑这么说,它拥有数十亿个参数,因此也是十亿维度的。但这个事实是否能转化为在高维空间中思考,我并不认为人类是这样,我也不认为这对大型语言模型同样适用。
Original English
Adam Brown: I don't think there's any sense I mean in which large language models naturally think in higher Dimensions more than humans do you could say well this large language models have billions of parameters that's like a billion dimensional space but you could say the same about the human brain that it has all of these billions of parameters and is therefore billion dimensional whether that that fact translates into thinking in uh billions of spatial Dimensions I don't really I don't really see that in the human and I don't think that applies to an LM either
Adam Brown: 是的,我想你可以想象,如果你看到了数百万个依赖于这种奇怪的张量数学的问题,那么就像人类通过训练来建立更好的直觉一样,AI也会发生同样的事情,它会看到更多的问题,并发展出对这些奇怪几何形状的更好表征。
Original English
Adam Brown: yeah I guess you could imagine that um you know if you just seen like a million different problems that rely on uh doing this uh weird tensor math then in the same way that maybe even a human gets trained up through that to build better intuitions the same thing would happen to AI just sees more problems and develop better representations of these kinds of weird geometries or something
Adam Brown: 我认为这绝对是真的,它确实看到了比我们任何人都多的例子,并且可能将发展出比我们更复杂的表征。
Original English
Adam Brown: I think that's certainly true that you know it it is definitely seeing more examples than any of us will'll ever see in our life and it is perhaps going to build more sophisticated representations than we have
Adam Brown: 是的,在物理学史上,一个突破往往在于你如何思考它,你采取哪种表征方式。有时人们开玩笑说,爱因斯坦对物理学最大的贡献是他发明的一种记法,叫做爱因斯坦求和约定(Einstein summation convention),它能让你更轻松地表达和思考这些事物,以更紧凑的方式,剥离掉其他一些东西。彭罗斯(Penrose)的一项伟大贡献是发明了一种新的记法来思考一些时空及其运作方式,这使得其他一些事物变得清晰。因此,显然,提出正确的表征一直是物理学史上一个极其强大的工具,也带来了许多极其重大的发展,这在某种程度上类似于在一些更应用的科学领域提出一种新的实验技术。
Original English
Adam Brown: yeah often in the history of physics a breakthrough is just you know how you think about it what representation you do it is sometimes jokingly said that Einstein's greatest contribution to physics was his uh a certain notation he invented called the Einstein summation convention which allowed you to more uh easily express and think about these things in a more compact way that strips strips away some of the other things you Penrose one of his great contributions was just writing down um a inventing a new notation for thinking about uh some of these space times and how they work that made certain other things clear so clearly coming up with the right repres presentation has been an incredibly powerful tool in the history of physics and and many incredibly large developments somewhat analogous to coming up with a new experimental technique in some of the more applied physic uh applied scientific domains
Adam Brown: 而且,我们确实希望,随着这些大型语言模型越来越好,它们能提出更好的表征,至少是它们自己的更好表征,但这可能与对我们来说好的表征不同。所以这里有一个有趣的问题。
Original English
Adam Brown: and yeah one would hope that uh as these large language models get better they come up with better representations at least better representations for them that may not be the same as a good representation for us so there's a there's an interesting question here
Adam Brown: 显然,这些模型知道很多,事实证明,即使是专业的物理学家也能提出他们不太熟悉的领域(fields)的问题并从中学习。但这是否提出了这样一个问题:我们认为它们很聪明而且越来越聪明,如果一个相当聪明的人记住了几乎每一个领域,了解开放性问题,了解其他领域的开放性问题以及它们可能如何与本领域联系,了解潜在的差异和联系,那么你可能会期望他们能够做出……不是爱因斯坦那样的概念性飞跃,但有很多事情,比如“嘿,镁(magnesium)与大脑中的这种现象相关,这种现象与头痛相关,因此也许镁补充剂可以治愈头痛”。诸如此类的基本联系,你自然会期待他们能做到。
Original English
Adam Brown: clearly these models know a lot and that's evidenced by the fact that even professional physicist can ask and learn about FS that they're less familiar with but this um doesn't this raise the question of we think these things are smart and getting smarter if a human that is reasonably smart had memorized uh basically every single field and knew about the open problems knew about the open problems in other fields and how they might connect to this field knew about um potential discrepancies and uh connections what you might expect them to be able to do is um not like Einstein level conceptual leaps but there are a lot of things where just like hey mag uh magnesium correlates with this kind of phenomenon in the brain this kind of Phenomenon correlates with headaches therefore maybe magnesium supplements cure headaches um these kinds of like basic connections you would anyways
Adam Brown: 这是否表明,就智能而言,大型语言模型甚至比我们预期的要弱?考虑到它们在知识方面拥有压倒性的优势,却还不能将这些转化为新的发现?
Original English
Adam Brown: does this suggest that llms are as far as intelligence goes even weaker than we might expect given the fact that given their overwhelming advantages in terms of knowledge they're not able to um already translate that into new discoveries
Adam Brown: 是的,它们肯定有人类所没有的优点和缺点。显然,它们的优点之一是它们阅读的量比任何人类一生中能读到的都多。我想,也许,再次以国际象棋程序为例可能是一个好比喻:它们经常考虑比任何人类棋手都多的可能局面(那是蒙特卡洛搜索),即使在人类水平的强度下,如果固定人类水平的强度,它们仍然会进行更多的搜索。所以,它们的评估能力可能不像人类那样自然。我认为物理学方面也会如此。
Original English
Adam Brown: yes they definitely have different strengths and weaknesses than humans and obviously one of their strengths is that they have read way more than any human will ever read in their entire life um I think maybe again the analogy with chess programs is is a good one here they will often consider way more possible positions that's Monte research than any human chess player ever would and yet they even at human level strength they're if you fix human level strength they're still doing way more search so that their ability to evaluate is maybe not quite as natural as a human so the same I think would be true of physics
Adam Brown: 如果有一个人类,读过的书和记住的东西和它们一样多,你可能会期望他甚至更强。斯科特·阿伦森(Scott Aaronson)最近,或者说一年多前,发文提到 GPT-4 在他的一门计算入门课(Introduction to Computing)上获得了 B 或 A- 的成绩,这肯定比我当年获得的分数高,所以我已经低于平均水平了。
Original English
Adam Brown: uh if you had a human who had had read as much and retained as much as as they had you might expect them to be even stronger Scott arenson recently or it was a year or so ago posted about the fact that the uh gbd4 got like a b or an A minus or something on his insured Computing class which is definitely a higher grade than I got and so I'm already below the waterline
Adam Brown: 嗯,但,是的,你知道,你在斯坦福讲授包括广义相对论(GR)在内的许多科目。我猜你一定用这些课程的考题去查询过这些模型,它们的表现是如何随时间变化的?
Original English
Adam Brown: um but uh yeah you know you teach a bunch of subjects including gr at Stanford um I assume you've been querying these models with questions from these exams how how how has their performance changed over time
Adam Brown: 是的,我……我拿了一个我几年前在斯坦福研究生生成性(generativity)课程上出的考试,然后给了这些模型。这相当非凡。三年前,零分,零分。一年前,它们表现还不错,可能算个弱学生,在平均水平之下,但现在它们基本上能通过考试。
Original English
Adam Brown: yeah I um I take an exam I gave years ago in my graduate generativity class at Stanford and give it to these models and it's pretty extraordinary three years ago zero zero uh a year ago they were doing pretty well maybe a weak student but in the distribution and now they essentially Ace the test
Adam Brown: 事实上,我退休了,那只是我自己的一个私人评估,我没有,你知道,它没有在任何地方发表,但我只是,我只是给它们这个东西,只是为了跟踪它们的表现,而且它的表现相当,相当强劲。它们,你知道,可能按照研究生课程的标准来说是简单的,但一门研究生课程,是的,关于广义相对论,它们在期末考试中几乎答对了所有问题。这只是在最近几个月,这些模型才做到这一点。
Original English
Adam Brown: in fact that that I'm retiring that that's just my own little private eval I don't you know it's not not published anywhere but I just I just give them this thing just to follow along how they're doing and it's pretty it's pretty strong um they you know may be easy by the standard of graduate courses but a graduate course yeah uh in general relativity and they get um pretty much everything right on the final exam that's just in the last couple of months that these these have been doing that
Adam Brown: 要通过一个测试,显然,它们可能读过所有的广义相对论教科书,但我认为要通过考试,你需要一些超越这些的东西。你是否会将物理学问题与数学问题区分开来?它们通常有两个组成部分:一是将应用题(word question)用你的物理学知识转化为数学问题,然后解决这个数学问题。这通常是这些问题的典型结构。所以你需要能够同时做到这两点:其中一部分是只有 LLM 才能做到的,而对其他东西来说并不那么容易,那就是第一步,将其转化为数学问题。
Original English
Adam Brown: what is required to a a test I obviously like they they probably have like read about all the generality textbooks but I assume to AC test you like need something be on that is there some you would characterize physics problems compared to math problems tend to have two components one is to sort of take this word question and like turn it using your physics knowledge into a a maths question yeah and then solve the maths question that's that tends to be the typical structure of these problems so you need to be able to do both the bit that's maybe you know only llms can do and wouldn't be so easy for other things is is step one of that is like turning into a math problem
Adam Brown: 我认为,如果你问它们一些困难的研究性问题,你肯定会遇到它们无法解决的问题,这毋庸置疑。但当我们试图开发对这些模型的评估方法时,我们可以注意到,就在几年前,尤其是在三年前,你可以从互联网上搜集到大量标准的高中数学题,而它们却做不出来。而现在,我们不得不聘请各领域的博士来构思问题,你知道,他们每天能想出一个绝妙的问题,或者类似的东西。随着这些 LLM 越来越强大,其评估它们的性能的难度也随之增加。
Original English
Adam Brown: um I think if you ask them hard research problems you certainly can come up with problems that they they can't solve that's that's for sure but it's pretty noticeable as we have tried to develop evaluations for these models that as recently as a couple of years ago certainly three years ago you just scrape from the internet any number of of problems that are standard totally standard high school math problems that they couldn't do and now we need to hire phds in whatever field and you know they they come up with one great problem a day or something you know the difficulty as these llms have got stronger the difficulty of evaluating their performance has has increased