AI智能的演进:预测的迷思与能力的飞跃 Dwarkesh Patel 2024-03-11

预测的局限性与智能的本质

演讲者回顾了AI的定标定律(Scaling Laws: 描述系统性能随规模增长而变化的规律)虽然可预测,但预测这些模型(Models: 指人工智能系统中用于处理和生成数据的算法结构)何时会爆发商业应用、或采取何种形式,则极其困难。他坦诚自己在这些方面的预测记录“糟透了”,并观察到普遍缺乏卓越的预测能力。尽管有时能准确预见某些方面,但理论上的展望常与现实脱节,能说对一成内容已属出类拔萃。由此引申出对智能(Intelligence)本质的思考,质疑了将智能视为从“村俗傻瓜”到爱因斯坦的简单线性尺度图示,并探讨这种抽象概念是否依然适用。演讲者指出,人类智能的范围极其广泛,且在不同任务上表现各异,暗示AI的发展或许不会遵循简单的线性进程。

Original English I feel like these scaling laws have been very predictable but then when you say like well you know when when is there going to be a commercial explosion in these models or what's the form it's going to be or are the models going to do things instead of humans or pairing with humans I feel like certainly my track record on predicting these things is is terrible but I also looking around I don't really see anyone who who track record is great I've been right about some things but I've still you know with these theoretical pictures ahead been wrong about most things being right about 10% of the stuff is you know sets you head and shoulders AB above um above above many people you know if you look back to I can't remember who it was kind of you know made these diagrams that are like you know here's here's the villageil idiot here's Einstein here's the scale of intelligence right and the V Village Idiot and Einstein are like very close to each other like that maybe that's still true in some abstract sense or something but it's it's not really what we're seeing is it we're seeing like that it seems like the human range is pretty Broad and doesn't we don't hit the human range in the same place or at the same time for different tasks right

AI能力边界:超人之处与局限

演讲者进一步阐述了AI模型能力的细微之处,指出它们在特定、受限任务上表现卓越,例如写一段不使用字母“e”的页面,甚至可能超越人类水平。然而,这与它们在证明相对简单的数学定理等领域则形成了鲜明对比,在这些领域,模型才刚起步,常犯“愚蠢的错误”,且缺乏广泛纠错或执行长任务的能力。这一观察引出了智能并非单一光谱的结论,而是由众多不同的专业领域和多样的技能组成,如记忆力,它们各自运作。演讲者推测,尽管智能在某种程度上可能存在于光谱上,但这个光谱本身是宽广且多维度的,这与早期的预期有所不同。

Original English like write write a sonnet you know in the style of cor MC McCarthy or something like I don't know I'm not very creative so I couldn't do that but like you know that's that's a pretty high level human skill right um and even the model is starting to get good at stuff of you know like constrained writing you know there's like write a you know write a page without using the letter e or something write a page about X without using the letter e like I think the models might be like superhuman or close to superum at that um but when it comes to you know I yeah I don't know prove relatively simple mathematical theorems like they're they're just starting to do the beginning of it they make really dumb mistakes sometimes and they they really lack any kind of broad like you know correcting your errors or doing some extended task and so I don't know it turns out that intelligence isn't isn't a spectrum there are a bunch of different areas of domain expertise there are a bunch of different like kinds of skills like memory is different I mean it's all it's all formed in the blob it's not it's all formed in the blob it's not complicated but to the extent it even is on the Spectrum the spectrum is also wide if you asked me 10 years ago that's not what I would have expected at all but uh I think that's very much the way it's turned out

AI发展意外:认知能力的解耦

AI发展过程中一个令人惊讶的方面是,认知能力似乎不像过去认为的那样相互关联。相反,模型似乎在不同时间学习不同的事物;例如,一个模型可能在编程方面表现出色,但在证明质数定理方面却遇到困难。这与早期认为不同认知能力会紧密联系、拥有单一潜在原理的预期形成了对比。演讲者指出,“智能”的概念本身似乎正在消解为一个连续统,促使人们将视角从抽象智能转向可观察的能力。回溯到2018年,演讲者绝不会预测到如今模型能以莎士比亚的风格写定理,或在带有开放性问题的标准化测试中表现出色,但这些模型显然不是通用人工智能(Artificial General Intelligence: 具备人类所有认知能力的AI)或达到人类水平。这种在基准测试中表现出的超乎寻常的性能与更广泛的人类能力现实之间的差距,仍然是一个显著的意外之处。

Original English one thing that's been surprising is like I thought things might click into place a little more than they do like you know I thought like different cognitive abilities might all be connected and there was more of one secret behind them but it's it's like the model just learns various things at different times you know and it can be like very good at coding but like you know it can't it can't quite you know prove the prime number theorem yet and I don't I mean I guess it's a little bit the same for for humans although it's it's weird the just deposition of things that can do and not I guess the main lesson is like having Theories of Intelligence or how intelligence works works like a lot of these words just just kind of like dissolve into a Continuum right they they just kind of like dematerialize I think less in terms of intelligence and More in terms of what what we see in front of us if you told me in 2018 we'll have models in 2023 like law to that can write theorems in the style of Shakespeare whatever theorem you want you want they can a standardized test with open-ended questions you know um just all kinds of really impressive things you would have said at that time I would have said oh you have AGI you clearly have something that is a human level intelligence where these while these things are impressive it clearly seems we're not at human level at least in the current generation and potentially for generations to come what explains discrepancy between super impressive performance in these benchmarks and in just like the things you could describe versus yeah generally so that that was one area where actually I was not presses and I was surprised as well

AI的演化路径:规模与目标的权衡

回顾早期对GPT-3等模型以及Anthropic公司初创时期的产品,演讲者回忆起一种“已掌握语言本质”的感觉,并质疑了进一步扩展规模的必要性,而非转向强化学习(Reinforcement Learning: 一种通过试错学习来优化决策的机器学习方法)等其他目标。最初的想法是,或许仅靠扩展就能持续带来改进,但这种想法已演变为对进一步扩展更有效,还是整合其他学习范式更有益的考量。模型在物理尺寸上比人脑(按突触 Synapses: 神经元之间传递信号的连接点,生物学上常用于比喻AI中的连接)衡量要小,但它们接受的训练数据却远超人脑——比人类18岁前接触的数据量还要多好几个数量级。这种差异导致了对生物学类比的怀疑,因为更小的模型却需要更多数据来执行复杂的人类任务,这其中存在悖论。演讲者认为,也许需要新的效率理解或不同的方法,但最终,重点正从依赖生物学比较转向直接衡量模型的能力。

Original English yeah um so when I first looked at gpt3 and you know more more so the kind of things that we built in the early days at at anthropic my my general sense was I you know I looked at these and I'm like it seems like they they've really grasped the essence of language I'm not sure how much we need to scale them up like maybe we maybe what's what's more needed from here is like RL and all and kind and kind of all the other stuff like we might be kind of near the you know I thought in 2020 like we can scale this a bunch more but I wonder if it's more efficient to scale it more or to start adding on these other objectives like like RL I thought maybe if you do as much RL as you know as as you've done pre-training for a for a you know 2020 style model that that's that's the way to go and scaling it up will keep working but you know is that is that really the best path and I I think it I don't know it just keeps going like I thought it had understood a lot of the essence of language but then you know there's there's kind of there's kind of further to go the models are maybe two to three orders a magnitude smaller than the human brain If you compare to the number of synapses while at the same time being trained on you know three to four or more orders of magnitude of data if you compare to you know number of words human a human sees as they're developing to age 18 we have to admit that that's a weird thing that doesn't match up and you know it's one reason I'm a bit you know skeptical of kind of biological analogies I thought in terms of them like five or six years ago but now that we actually have these models in front of us as artifacts it feels like almost all the evidence from that has been screened off by what we've seen and what we've seen are models that are much smaller than the human brain and yet yet can do a lot of the things that humans can do and yet paradoxically require a lot more data um so maybe we'll discover something that makes it all efficient or maybe we'll understand why the discrepancy is present but at the end of the day I don't think it matters right if we keep scaling the way we are I think what's more relevant at this point is just measuring the abilities of the model and seeing how far they are from humans and they don't seem terribly far to me

知识汇聚与创造力:AI的未来展望

演讲者回应了一个关键问题:模型拥有整个人类知识语料库(Corpus: 大规模的文本或数据集合,用于训练AI模型),为何不像智能程度中等的人类那样,能做出单一新连接并引发现象级发现?他承认这一差距,但认为模型确实展现出“普通创造力”,例如以Cormac McCarthy或Barbie的风格写十四行诗,这涉及到普通人也会进行的全新连接。虽然当前模型尚未产生重大科学发现,但演讲者相信随着持续扩展,模型技能水平的提高将带来改变。AI模型的一个显著优势在于其比人类掌握的知识更多。这在生物学等复杂领域尤为重要,因为发现的前提是需要海量的知识储备。演讲者总结道,当前模型已积累了大量知识,并正处于能够综合这些信息并取得重要突破的边缘。

Original English what do you make of the fact that these things have basically the entire Corpus of human knowledge memorized and as far as I'm aware they haven't been able to make like a single new connection that has led to a discovery whereas if even a moderately intelligent person had this much stuff memorized they noticed oh this thing causes this symptom this other thing also causes this symptom you know there's a medical cure right here right what shouldn't we be expecting that kind of stuff I'm not I'm not sure I mean I think you know I don't know these words Discovery creativity like it's one of the lessons I've learned is that in in you know in kind of the Big Blob of compute often these these ideas often end up being kind of fuzzy and Elusive and hard to track down but I think I think there is something here which is I think the models do display a kind of ordinary creativity again again you know the kind of like you know write a write a Sonet you know in the style of cormic McCarthy or Barbie or you know like there is some creativity to that and I think they do draw you know new connections of the kind that an ordinary person would draw I I agree with you that there haven't been any kind of like I don't know like I would say like big scientific discoveries I think that's a mix of like just the model skill level is not is not high enough yet that I think is going to change with the with the scaling I do think there's an interesting point about well the models have an advantage which is they know a lot more than us you know like should should they have an advantage already even even if they skill level isn't isn't isn't quite High maybe that's kind of what you're getting at I don't really have an answer to that I mean it seems certainly like memorization in facts and drawing connections is an area where the models are ahead and I I I do think maybe you need those connections and you need a fairly high level of skill I do think particularly in the area of biology for better and For Worse the complexity of biology is such that the current models know a lot of things right now and that's what that's what you need to make discoveries and draw it's not like physics where you need to you know you need to think and come up with a formula in biology you need to know a lot of things right and so I do think the models know a lot of things and they have a skill level that's not quite high enough to put them together and I think they are they are just on the cusp of being able to put these things together
📌 文中提及的人物和组织

公司/组织: Anthropic

产品/模型: GPT-3

关键字: intelligence scaling-laws ai-models prediction cognitive-abilities creativity