生成式AI在科技史上的位置:一个反主流观点 Internet of Bugs 2025-11-03

生成式AI的炒作与历史地位

大约一年前,我制作了一个关于AI炒作失控的视频。

View/Hide Original English

I did a video a year or so ago about how AI hype was out of control.

我很遗憾地告诉大家,情况并没有好转。
View/Hide Original English

I regret to inform you, it hasn't gotten any better.

我那个关于炒作的视频主要关注了AI行业如何利用演示和新闻稿来推销其产品夸大的承诺,而这些承诺实际上无法兑现。
View/Hide Original English

So that video I did on hype focused on how the AI industry use demos and press releases to try to sell exaggerated promises about their products that they can actually deliver.

本视频则侧重于AI行业如何试图销售一种被简化了的历史版本,以使当前一代的AI技术在历史中占据比实际更重要的位置。
View/Hide Original English

This video focuses on how the AI industry tries to sell a trivialized version of the past to make the current generation of AI technology seem to have a more important place in history than it actually does.

这是萨姆·奥特曼(**Sam Altman**: OpenAI首席执行官)与《纽约时报》(**The New York Times**: 美国著名报纸)的对话:
View/Hide Original English

So here's Sam Altman talking to the New York Times:

每个人对AI都有自己的类比,你知道,我看到桑达尔(**Sundar Pichai**: Google首席执行官)也会在这里。
View/Hide Original English

Everybody has their analogy for what AI is like, you know, I saw Sundar is going to be here.

他把它比作电力。
View/Hide Original English

He talks about it like electricity.

很多人把它比作工业革命(**Industrial Revolution**: 18世纪末至19世纪初的工业技术变革)。
View/Hide Original English

A whole bunch of other people talk about like the industrial revolution.

有些人把它比作文艺复兴(**Renaissance**: 14至17世纪欧洲文化艺术复兴)。
View/Hide Original English

Some people talk about like the Renaissance.

我喜欢的一个是晶体管(**Transistor**: 一种半导体器件,是现代电子技术的基础)。
View/Hide Original English

The one I like is the transistor.

对我来说,将**生成式AI**(Generative AI: 能够生成文本、图像等新内容的AI)与电力、工业革命、文艺复兴或晶体管相提并论,显然是荒谬的。
View/Hide Original English

To me, comparing generative AI to electricity, the industrial revolution, the renaissance or the transistor is just obviously ridiculous.

但这种说法细节太少,而且这些事件发生得太久远,以至于很难进行知情的讨论,这意味着这些主张在很大程度上未受挑战,当然,也被记者和影响者们无休止地、不加批判地重复。
View/Hide Original English

But there are so few details in that kind of statement and the events are so long ago that it's difficult to even have an informed discussion about it, which means these claims go largely unchallenged and of course, repeated endlessly and uncritically by reporters and influencers.

这类声明的主要问题在于,AI(即人工智能)的定义非常模糊,而且是故意为之,这使得反驳它变得非常困难。
View/Hide Original English

The primary problem with statements like these is that the definition of AI, e.g. artificial intelligence, is very ambiguous and intentionally so, which makes it really difficult to argue with.

无论AI在100年、1000年或10000年后可能变成什么样,它很可能像某些基础技术一样具有影响力。
View/Hide Original English

Whatever AI might become in 100 years or 1000 years or 10,000 years may well be as impactful as some fundamental technology.

另一方面,OpenAI(**OpenAI**: 人工智能研究实验室)在资金耗尽之前能完成什么,则完全是另一回事。
View/Hide Original English

On the other hand, what OpenAI can get done before their money runs out is an entirely different thing.

这就是为什么很难反驳这些观点的原因之一,因为他们对“AI”一词的含义可以故意地逐句改变。
View/Hide Original English

That's one of the reasons why it's so hard to argue with these points because what they mean by the term "AI" can change from sentence to sentence deliberately.

所以今天,我将在一个更具体的声明背景下讨论AI,这个声明是AI行业中另一个人提出的,关于AI与几十年来软件其他进步的比较。
View/Hide Original English

So instead today, I'm going to discuss AI in the context of a much more concrete statement made by someone else in the AI industry, a statement about how AI compares to other advances in software over the decades.

因为如果我能将AI置于过去几十年科技行业的背景下,并且如果这能帮助你相信AI的影响力不如晶体管所实现的几项其他软件创新,那么这有望很好地解释为什么我认为将AI与晶体管,甚至更不可能的电力或工业革命相提并论,简直是胡说八道。
View/Hide Original English

Because if I can put AI in the context of the last few decades of the tech industry, and if that helps convince you that AI is less impactful than several other software innovations that were enabled by transistors, that hopefully will go a long way to explain why I think that comparing AI to something like transistors or even less possibly electricity or the industrial evolution is just a load of Bullsh-----

欢迎来到“Internet of Bugs”。
View/Hide Original English

Welcome to the Internet of Bugs.

我叫卡尔(**Carl**: 本视频主持人)。
View/Hide Original English

My name is Carl.

我从事软件专业工作已经超过35年了,我在YouTube上谈论软件也有一年半左右了,我正在尽自己的一份力,让互联网变得不那么“buggy”,这包括试图向人们解释为什么相信**生成式AI**的炒作是一个非常非常糟糕的主意。
View/Hide Original English

I've been a software professional for more than 35 years now, and I've been talking about software here on YouTube for a year and a half or so, and I'm trying to do my part to make the Internet a less Buggy place, which includes trying to explain to people by believing the hype about generative AI is a really, really bad idea.

重新定义AI的讨论范围与时间轴

现在,我不再试图反驳“AI就像晶体管”这种模糊的说法,而是要讨论一个不那么模糊、更谦逊的声明,它来自最近一次Y Combinator(Y Combinator: 知名创业加速器)会议上的一次演讲。

View/Hide Original English

Now, instead of trying to argue with a vague statement like "AI is like a transistor," I'm going to discuss a much less vague, much more modest statement from a presentation given at a recent conference for the startup accelerator Y Combinator.

这是特斯拉(**Tesla**: 知名电动汽车和清洁能源公司)前AI负责人演讲中的引述:
View/Hide Original English

So here's that quote from the presentation by the former head of AI at Tesla:

“我认为,大致来说,软件在如此根本的层面上已经70年没有太大变化了,然后我认为在过去几年里,它迅速地改变了大约两次。”
View/Hide Original English

"I think, roughly speaking, software has not changed much on such a fundamental level for 70 years, and then it's changed, I think, about twice, quite rapidly in the last few years."

现在,我想明确一点,我并不是要专门攻击这位演讲者。
View/Hide Original English

Now, I want to be clear that I'm not trying to attack this speaker specifically.

他今天所做的演讲,与我前面引用的奥特曼的言论相比,已经相当温和了,但这确实是AI领域知名人士经常发表的言论的一个例子,这些言论让公众认为AI比实际情况重要得多。
View/Hide Original English

The presentation he gave that I'm going to talk about today was pretty tame compared to Altman's quote I gave earlier, but it is an example of the kind of statements that are regularly made by AI luminaries that make the general public think that AI is a much bigger deal than it actually is.

而这反过来又会误导那些相信炒作的公众做出他们和社会以后可能会后悔的糟糕决定。
View/Hide Original English

And that in turn, that tricks the hype believing public into making poor decisions that they and society may well regret later.

我们中的一些人已经后悔了。
View/Hide Original English

That some of us already regret.

我选择这个演讲作为切入点,是因为这个声明有一个更明确的时间范围:70年。
View/Hide Original English

I'm picking this presentation as a jumping off point because this statement has a much more definite timeframe, 70 years.

它没有使用模糊的AI营销术语,这个术语可以根据演讲者的意愿在每句话中改变含义。
View/Hide Original English

It doesn't use the vague AI marketing term, which can mean whatever the speaker wants to this sentence.

相反,它谈论的是过去几年发生的变化,这意味着与此引述相关的AI以及我将在本视频中讨论的AI是**生成式AI**,它也被称为**大型语言模型**(Large Language Models, LLMs: 能够理解和生成人类语言的AI模型)。
View/Hide Original English

Instead, it's about changes that have happened in the last few years, which means that the AI relevant to this quote and the AI I'll be discussing in this video is generative AI, which is also called large language models.

此外,由于这个引述是关于软件的,它更适合我的频道、我的受众和我的经验。
View/Hide Original English

Also, since this quote is about software, it's much better fit from my channel and my audience and my experience.

尽管我才五十多岁,没有经历过这70年,但我已经经历了大部分时间,并且编写软件的时间超过了一半。
View/Hide Original English

And although I'm only in my fifties and I haven't lived through the last 70 years, I have lived through most of it and I've been writing software for more than half that time.

所以,我觉得自己有资格谈论它。
View/Hide Original English

So it's something I'm feel qualified to talk about.

所以今天我将提出一个观点,即软件不仅在过去70年里在根本层面上发生了几次巨大的变化,而且我们现在正在经历的以及过去几年所看到的变化,相比之下,实际上并没有那么根本。
View/Hide Original English

So today I'm going to make a case that not only has software changed drastically at a fundamental level several times over the last 70 years, but that the changes we are living through now and what we've seen over the last few years are by comparison, not actually even all that fundamental.

我甚至会尝试将AI置于我认为它在技术史上的应有位置,以及在我看来它真正的含义,但我们稍后会详细讨论。
View/Hide Original English

I'm even going to try to put AI in the context of where I think it ought to be in the history of technology and what it really means in my opinion, but we'll get to more of that later.

所以70年前大约是1955年中期,但我不想担心意外地谈论70年半前发生的事情。
View/Hide Original English

So 70 years ago was about mid 1955, but I don't want to have to worry about accidentally talking about something that happened 70 and a half years ago.

所以我将我的起始日期定为1956年。
View/Hide Original English

So I'm going to start my date at 1956.

对于“过去几年”的截止日期,我将选择2015年。
View/Hide Original English

For my cutoff date for the last few years, I'm going to pick 2015.

同样,我将尽量避免部分年份。
View/Hide Original English

Again, I'm going to try to avoid partial years.

所以我将说从1956年1月1日到2015年12月31日,这段时间完全符合“软件 supposedly 没有太大变化”的70年之内。
View/Hide Original English

So I'm going to say that from January 1st, 1956 to December 31st, 2015 is well within the 70 years that supposedly software has not changed much.

而从2016年1月1日到现在这段时期,应该肯定涵盖了“过去几年”,据称在这期间软件迅速改变了两次。
View/Hide Original English

And the period from January 1st, 2016 to now should definitely cover the last few years, supposedly during which it's changed twice quite rapidly.

2016年1月大约是“Attention Is All You Need”(**Attention Is All You Need**: 一篇2017年发布的开创性论文,引入了Transformer架构)论文发表前一年半,这篇论文开启了最新一轮的AI热潮。
View/Hide Original English

January of 2016 is roughly a year and a half before the "attention is all you need" paper that kicked off the latest general of AI craze.

所以我认为这是一个宽泛的截止日期。
View/Hide Original English

So I think it's a generous cutoff date.

我将提出一个观点,即与之前大约60年的变化相比,2016年以来的变化坦白说相当无聊。
View/Hide Original English

And I'm going to make the case that the changes since 2016 have honestly been pretty boring in comparison to the changes from the 60ish years before that.

我将提出一个观点,即从上下文来看,**大型语言模型**与前一个时代的许多变化相比,实际上并没有那么根本。
View/Hide Original English

I'm going to make the case that, in context, large language models aren't actually all that fundamental compared to many of the changes from the previous era.

1956年的软件状态:从打孔卡到键盘输入

那么第一步,让我们谈谈1956年的软件状态。

View/Hide Original English

So step one, let's talk about the state of software in 1956.

1956年,人们正在进行实验,试图找出如何让计算机从键盘而不是当时使用的打孔卡接收输入。
View/Hide Original English

In 1956, experiments were being done trying to figure out how to get a computer to take input from a keyboard instead of the punch cards that were used at the time.

1956年是第一本**Fortran**(Fortran: 第一批高级编程语言之一,主要用于科学计算)参考手册出版的年份,它是第一种第三代语言或编译语言,尽管该语言还需要一段时间才能稳定下来。
View/Hide Original English

1956 was the year the first reference manuals published for the first third generation language or compiled language Fortran, although it would be a while before the language actually stabilized.

**Algol**(Algol: 一种早期的算法语言)将在两年后出现。
View/Hide Original English

Algol was two years in the future.

**Cobol**(Cobol: 面向商业的通用语言)将在三年后出现。
View/Hide Original English

Cobol was three years in the future.

美国第一个计算机编程新生课程将在两年后在**卡内基梅隆大学**(Carnegie Mellon: 美国著名研究型大学)开始。
View/Hide Original English

The first freshman computer programming course in the US would begin at Carnegie Mellon in two years.

而第一个支持软件开发的通用分时系统,一个学生可以实际完成作业的系统,将在1956年后的五年在**麻省理工学院**(MIT: 美国著名研究型大学)发布。
View/Hide Original English

And the first general purpose time sharing system that supported software development, something that students could actually do their homework on, would be released at MIT five years from 1956.

现在,我不知道你怎么样,但我会认为能够向计算机输入文字对于开发软件来说是相当基础的。
View/Hide Original English

Now, I don't know about you, but I would consider the ability to type into a computer to be pretty fundamental to developing software.

我根本没有足够的耐心来处理打孔卡。
View/Hide Original English

I don't have anywhere near enough patience to deal with punch cards.

这意味着我会说,在过去70年里已经发生了一件根本性的事情。
View/Hide Original English

So that means I'd say there's already one fundamental thing that's happened in the last 70 years.

此外,我认为有史以来第一堂大学编程课程的引入对于开发软件来说是相当基础的。
View/Hide Original English

Additionally, I think the introduction of the first ever college class and how to program is pretty fundamental developing software.

而且我认为,作为学生,实际拥有一台可以练习编程的计算机对于开发软件来说是相当基础的。
View/Hide Original English

And I think actually having a computer that exists that you can practice programming on as a student is pretty fundamental to developing software.

但让我带你了解更多,因为我认为这将有助于说明我们当前的**生成式AI**在历史上真正所处的位置。
View/Hide Original English

But let me take you through more of it, because I think it will help illustrate where our current generative AI really fits into history.

我将跳过大量内容。
View/Hide Original English

I'm going to skip a ton of stuff.

我不会严格按时间顺序进行。
View/Hide Original English

I'm not going to go strictly chronologically.

我将很快地浏览这些内容,并且为了节省时间,我将进行很多过度简化,但我希望你能从中看到这些联系,这会让你觉得有趣。
View/Hide Original English

I'm going to go through these pretty quickly, and I'm going to oversimplify a lot for the sake of time, but I'm hoping it will be interesting to you to see the connections.

1970年代:数据处理与第三代语言的兴起

我们的下一站是1970年代。

View/Hide Original English

Our next stop is the 1970s.

这样我就可以向你介绍一个我们现在不怎么谈论的旧概念,叫做**数据处理**(Data Processing: 对数据进行收集、存储、检索、转换和分析的过程)。
View/Hide Original English

So I can introduce you to an old concept. We don't talk about much anymore, which is called data processing back when it was all about punch cards.

当一切都围绕着打孔卡时,它并不是我们现在所认为的编程。
View/Hide Original English

It wasn't programming in the way we think of it now, really.

除了错误处理,它主要只是数学运算。
View/Hide Original English

It was mostly just mathematical operations, except for the error handling.

你可以想象,当你从一堆由没有屏幕的人手工打孔的物理卡片中读取数据时,打字错误是不可避免的,卡片的折叠和撕裂、传感器中的灰尘等等也是如此。
View/Hide Original English

As you can imagine, when you are reading data of a bunch of physical cards that are punched by hand by someone who didn't have a screen, typos were inevitable, and so were folds and tears and cards, dust in the sensors, et cetera, et cetera.

所以这类指令中最复杂的部分是确保你不会因为错误数据而得到糟糕的结果或崩溃。
View/Hide Original English

So the most complicated part of the instructions for this kind of thing was making sure that you didn't get bad results or crashes from bad data.

我们将继续回顾**数据处理**这个概念,因为我们浏览这个列表,它几乎就像与软件开发并行发生的不同轨迹,原因将在最后变得清晰。
View/Hide Original English

We're going to keep revisiting the idea of data processing as we go through this list, almost like a different track happening in parallel with software development for reasons that will become clear at the end.

我们接下来要讨论的软件开发方面的进展是**C语言**(C: 一种高级编程语言)等**第三代语言**(Third Generation Languages: 相比机器语言和汇编语言更接近人类语言的编程语言)的兴起。
View/Hide Original English

The next development in the software development side we're going to talk about was the rise of third generation languages like C.

**编译器**(Compiler: 将高级语言代码转换为机器码的程序)将更少、更简单、更易读的指令转换为更多、更低级、高度详细的计算机可以理解的指令。
View/Hide Original English

The compiler turned fewer, much simpler and easier to read instructions and many more lower level, highly detailed instructions that the computer could understand.

这被期望能提高程序员的生产力,而且确实如此。
View/Hide Original English

This was expected to make programmers more productive, and it did.

有些人认为这意味着需要更少的程序员。
View/Hide Original English

And some people thought that wouldn't mean fewer programmers would be needed.

但新语言所带来的额外能力意味着软件可以做更多的事情。
View/Hide Original English

But the additional capabilities that the new language is enabled meant software could do more.

所以我们实际上最终有了更多的程序员。
View/Hide Original English

And so we actually ended up with more programmers.

我们姑且称之为我们稍后会再讨论的主题。
View/Hide Original English

Let's just say that's a theme we'll come back to you later.

关系型数据库:现代生活的基石

回到数据处理领域,关系型数据库(Relational Databases: 基于关系模型组织和存储数据的数据库)被发明了。

View/Hide Original English

Back in the data processing land, relational databases were invented.

早期有多种查询语言。
View/Hide Original English

Early on, there were multiple query languages.

我学的第一种实际上不是**SQL**(SQL: 结构化查询语言,用于管理关系型数据库)。
View/Hide Original English

I'm actually the first one I learned wasn't SQL.

但现在几乎所有人都已标准化为**SQL**。
View/Hide Original English

But now everyone is pretty much standardized on SQL.

我无法夸大这些数据库对现代生活的改变有多大。
View/Hide Original English

I can't overstate how much these databases change modern life.

世界上绝大多数可查询的数据,在观看此视频的每个人生命的大部分时间里,都以这种数据库形式保存。
View/Hide Original English

The vast majority of the queryable data in the entire world has been kept in this kind of database for the majority of the lives of everyone watching this.

如果你深入研究现代世界的许多方面,尤其是任何与金钱有关的事情,你很可能会发现**SQL数据库**是其核心。
View/Hide Original English

If you drill down on so much of the modern world, especially anything that has anything to do with money, you'll likely find a SQL database at the core of it.

所有这些数据库中一个非常有用的概念叫做**约束**(Constraints: 数据库中用于验证数据完整性的规则)。
View/Hide Original English

One of the really useful concepts in all of these databases is called constraints.

它是一种在数据写入数据库之前验证数据的方法,这大大加快了许多查询和错误检查的速度,比打孔卡时代快得多。
View/Hide Original English

It's a way of validating data before it's written to the database, which drastically speeds up a lot of the querying and error checking than in the punch card days from a programming perspective.

从编程角度来看,**关系型数据库**允许将以前通过手工或专用制表机完成的**数据处理**任务与新的第三代编程语言相结合,为无数以前不可能的现实世界任务开辟了无限的可能性。
View/Hide Original English

Relational databases allowed the intersection of the kind of data processing tasks that have been done by hand or by dedicated tabulating machines and the new third gen programming languages opening up a myriad of possibilities for uncountable numbers of real world tasks that were now possible that wouldn't have been before programming language wise.

面向对象编程与索引全文搜索

在编程语言方面,我们的下一个进步是面向对象编程(Object Oriented Programming, OOP: 一种编程范式,将数据和操作封装在对象中),当时它似乎是一个很棒的主意。

View/Hide Original English

Our next advances object oriented programming, which, well, seemed to be like a great idea at the time.

其中一些甚至确实如此。
View/Hide Original English

And then some of it even was.

同样,这里的目标是提高程序员的生产力,以便我们能更快、更少地完成更多工作。
View/Hide Original English

Again, the goal here was to make programmers more productive so we could do more faster with less.

**OOP**实现了许多没有它会更难完成的事情,稍后会有更多讨论。
View/Hide Original English

There were a lot of things at OO enabled that would have been much harder to do without it, more on that in a bit.

而且,它再次增加了对程序员的需求。
View/Hide Original English

And again, it just made more demand for programmers.

回到数据领域,我想强调**索引全文搜索**(Indexed Full Text Search: 通过索引快速搜索大量文本数据)。
View/Hide Original English

Back in data land, I want to highlight indexed full text search.

这些索引技术首次使得如此多的**人类知识总和**能够被快速访问。
View/Hide Original English

These indexing techniques allowed for the first time for so much of the sum of human knowledge to be accessed quickly.

你们中的大多数人现在可能认为这是理所当然的,但对于那些伴随图书馆卡片目录长大的人来说,我可以告诉你们,这是人们学习方式上一个根本性的改变。
View/Hide Original English

Most of you probably take this for granted now, but it's one of those people that grew up with card catalogs and libraries. I can tell you this was a fundamental change in the very process of how people learn.

但它也很重要,因为它帮助我说明了一个有趣的编程问题,那就是规模。
View/Hide Original English

But it's also significant in that it helps me illustrate an interesting programming problem that comes up here, which is scale.

到目前为止,理论上可以验证已编写的大多数编程任务的正确答案。
View/Hide Original English

Up until now, it's theoretically possible to verify the correct answer for most programming tasks that had been written to this point.

但**全文搜索**让你坚定地进入了处理大量数据的领域。
View/Hide Original English

But full tech search puts you firmly in the realm of processing huge stacks of data on punch cards.

你可以在小数据集上测试你的代码,并验证它在这些数据集上得到了正确的答案。
View/Hide Original English

You can test your code on small data sets and verify that for those, it gets the right answers.

但一旦你在大数据集上运行它,你就无法再验证它的正确性了。
View/Hide Original English

But once you turn it loose on the huge data sets, then you can't verify it's correct anymore.

例如,想象一下你正在对图书馆所有书籍进行**全文搜索**,寻找包含某个短语的内容。
View/Hide Original English

So for example, imagine you're doing full tech search across all the books in a library and you're looking for things that contain a certain phrase.

如果有一本书确实包含该短语,但由于某种原因它没有出现在搜索结果中,除非你偶然发现那本书中的那个短语,否则你不会知道你从计算机程序中得到的答案是不完整的。
View/Hide Original English

If there's one book that does contain that phrase, but it's missing from the search output, for whatever reason, unless you happen to stumble across that phrase in that book, you don't know that the answer you got from the computer program was incomplete.

图形用户界面、数据仓库与互联网时代

面向对象编程为现代图形用户界面(Graphical User Interfaces, GUIs: 允许用户通过图形图标和视觉指示器与电子设备交互的界面)铺平了道路。

View/Hide Original English

So object oriented programming led the way to modern graphical user interfaces.

我们以前也有一些,但创建具有大小属性等的窗口对象的能力使其更具可用性。
View/Hide Original English

We had some before, but the ability to make a Window object that has size properties, et cetera, et cetera made it a lot more usable.

随之而来的是**集成开发环境**(Integrated Development Environments, IDEs: 程序员用于软件开发的一套工具),结合一些搜索技术,我们得到了**自动补全**(Autocompletes: 自动完成代码或文本输入的功能)、**语法高亮**(Syntax Highlighting: 以不同颜色显示代码不同部分的功能),这非常有帮助。
View/Hide Original English

With this, we also got integrated development environments and combined with some of the search tech, we got autocompletes, syntax highlighting, which was incredibly helpful.

你们中的大多数人不知道,当你必须记住并正确拼写代码中的每一个东西才能使其编译时,编程速度会慢多少。
View/Hide Original English

Most of you have no idea how much slower you program when you have to remember and correctly spell every single thing in your code to get it to compile.

数据轨迹上的下一个是**数据仓库**(Data Warehousing: 用于存储大量历史数据以供分析的系统)。
View/Hide Original English

Next up on the data track is data warehousing.

将大量数据放入一个巨大的数据库中,然后事后找出你想要从中获取什么的想法。
View/Hide Original English

The idea of putting lots and lots of stuff at a giant database and then figuring out after the fact what you want to get out of it.

同样,错误处理和过滤掉垃圾数据是一个巨大的问题。
View/Hide Original English

Again, error handling and filtering out garbage data is a huge problem.

大量使用索引是使搜索时间可容忍的唯一方法。
View/Hide Original English

A heavy use of indexing is the only way to make search times tolerable.

一旦你达到一定量的数据,你就没有一个好的方法来验证你的代码是否仍然正常工作。
View/Hide Original English

And once you get to a certain amount of data, you don't have a good way of verifying if your code is still working correctly.

尽管如此,它们还是非常有用的。
View/Hide Original English

That said, they're incredibly useful.

我们稍后会回到**数据仓库**。
View/Hide Original English

We'll come back to data warehouses in a while.

现在,是时候谈论互联网了。
View/Hide Original English

Now, it's time to talk about the internet.

我可以在这里列出所有底层的基础技术,例如**TCP/IP**(TCP/IP: 互联网协议套件,是互联网通信的基础)、**生成树协议**(Spanning Tree: 一种网络协议,用于防止网络环路)、**CDMA**(CDMA: 码分多址,一种无线通信技术)、**BGP**(BGP: 边界网关协议,用于在互联网上路由数据)。
View/Hide Original English

I can make a whole list of the underlying fundamental pieces of tech here, TCP/IP, spanning tree, CDMA, BGP.

我也会把**加密**(Encryption: 将信息转换为代码以防止未经授权访问的过程)和**密钥交换**(Key Exchange: 在通信双方之间安全地共享加密密钥的方法)放在这里。
View/Hide Original English

I'm going to put encryption and key exchange here too.

但就软件开发而言,它所实现的一大软件工程进步是**依赖项**(Dependencies: 软件项目所需的外部代码或库)和**包管理器**(Package Managers: 自动化软件安装、升级、配置和删除的工具)。
View/Hide Original English

But for software development purposes, one of the big software engineering advances that enabled was dependencies and package managers.

我遇到的第一个是**CPAN**(CPAN: Perl语言的包管理器)。
View/Hide Original English

The first one I ran into was called CPAN.

它是为**Perl**(Perl: 一种高级脚本语言)语言设计的,然后是**Java**(Java: 一种广泛使用的编程语言)的**Ant/Maven**(Ant/Maven: Java项目的构建工具和包管理器),但你可能更熟悉**Python**的**pip**(Python's pip: Python的包安装程序)或**NodeJS**的**NPM**(NodeJS's NPM: Node.js的包管理器)之类的。
View/Hide Original English

It was for the Perl language, then Ant/Maven for Java, but you're likely more familiar with Python's pip or NodeJS's NPM or the like.

通过帮助开发者下载代码而不是自己编写,我们希望能够再次提高生产力。
View/Hide Original English

By helping developers download code instead of having to write it themselves, it was hoped that we could be more productive again.

我知道你们中的一些人现在看到AI从一个简单的提示生成大量代码而感到震惊。
View/Hide Original English

I know that some of you are freaked out now by watching AI generate a bunch of code from a simple prompt.

让我向你保证,从我们每个人都必须从头编写所有代码,到只需在配置文件中输入一个包名就能将数千行专用代码包含到你的项目中,行业所获得的生产力提升至少同样巨大,甚至更大。
View/Hide Original English

Let me assure you the productivity boost that the industry got by going from each of us having to write all of our own code from scratch, to including thousands of lines of purpose-built code into your project just by typing one package name into a config file, was at least as dramatic, if not more.

现在,这有点奇怪,但在数据领域,我想谈谈**有损压缩**(Lossy Compression: 一种数据压缩方法,通过丢弃部分数据来减小文件大小,但会损失一些信息)。
View/Hide Original English

Now, this thing is going to be weird, but next to the data space, I want to talk about lossy compression.

经典例子是**JPEG**(JPEGs: 一种图像压缩格式)、**MP3**(MP3s: 一种音频压缩格式)和**MP4**(MP4s: 一种视频压缩格式),音频和视频类型的东西。
View/Hide Original English

Classically, this is like JPEGs and MP3s and MP4s, audio and video type stuff.

基本上,它是一种让我们存储比原始数据少得多的数据的方法。
View/Hide Original English

Basically, it's way for us to store a lot less data than we started with.

然后当我们从存储中取出它时,我们会弥补我们没有实际存储的部分。
View/Hide Original English

And then when we pull it back out of storage, we make up the parts of it we didn't actually store.

所以它希望能看起来足够接近原始。
View/Hide Original English

So it hopefully looks close enough.

如果我们丢弃了太多信息,或者如果我们要压缩的原始数据具有某些特性,我们就会得到很多**压缩伪影**(Compression Artifacts: 有损压缩过程中丢失信息而产生的视觉或听觉失真),这有时取决于你想要做什么,有时又会破坏体验,取决于我们想要实现什么。
View/Hide Original English

If we throw away too much information or if the original we're trying to compress has certain characteristics, we'll get a lot of compression artifacts, which sometimes is good enough depending on what you're trying to do and sometimes destroys the experience, depending on what we're trying to accomplish.

但这不限于媒体。
View/Hide Original English

But it's not limited to media.

想想如何存储大量历史股票价格数据。
View/Hide Original English

Think about something like trying to store a lot of historical stock price data.

你可以每时间段有更少的样本,并且在存储之前对数值进行四舍五入。
View/Hide Original English

You can have fewer samples per time span and you can round the values before you shore them.

如果你只是想分析趋势,那可能就足够了。
View/Hide Original English

If you're just trying to analyze trends, that's probably good enough.

但重要的是你要了解它的局限性。
View/Hide Original English

But it's important that you understand its limitations.

然而,没有它,我们就无法存储或传输互联网上一直传输的数据量。
View/Hide Original English

Without this, though, we just would not be able to store or transmit the amount of data that crosses the internet all the time.

这意味着没有播客,没有YouTube,我们的互联网将是一个不那么有趣的地方。
View/Hide Original English

And it would mean no podcasts, no YouTube and our internet would be a much less interesting place.

所以这是一个大问题。
View/Hide Original English

So it's a big deal.

LLM的真实定位:有损压缩的数据处理技术

我还可以提到许多其他基础技术,比如无代码网站(No-code websites: 允许用户无需编写代码即可创建网站的平台),如WordPress(WordPress: 一种流行的内容管理系统)和Squarespace(Squarespace: 一种网站建设平台),用于分布式网络规模计算的Map-Reduce(Map-Reduce: 一种分布式计算编程模型),虚拟机(Virtual Machines: 模拟计算机系统的软件),云技术(Cloud Tech: 通过互联网提供计算资源的服务),它允许软件运行而不受其运行的物理硬件结构的限制。

View/Hide Original English

There are a bunch of other fundamental technologies I could have mentioned, like no-code websites, like WordPress and Squarespace, Map-Reduce for distributed web scale computing, virtual machines, cloud tech, that lets software run without being constrained by the structure of the physical hardware it's running on.

更不用说**智能手机**(Smartphones: 具有先进计算能力的移动电话)和**移动计算**(Mobile Computing: 通过便携式设备进行计算),但我们已经讲了足够多的内容,我可以为你提供**大型语言模型**的另一种定位。
View/Hide Original English

Not to mention smartphones and mobile computing, but we've gotten through enough that I can give you an alternative placement for the large language models.

现在,我认为**LLM**只是下一代**数据处理**技术,它使用**有损压缩**,对云规模索引进行搜索。
View/Hide Original English

Now, I would argue that LLMs are just the next generation of data processing tech using lossy compression, searching against a cloud scale index.

有些人称之为“思考”的,实际上只是带有非常丰富的非结构化查询语言的搜索,而所谓的“**幻觉**”(Hallucinations: AI生成不准确或虚构信息)实际上只是**压缩伪影**,它通过编造词语来填补空白,就像压缩的JPEG图像会弥补像素,或者压缩的视频会弥补图像帧的块一样。
View/Hide Original English

What some people call "thinking" is really just search with a very rich unstructured query language and the so-called "hallucinations" are effectively just compression artifacts, where it's making up words to fill in the gaps, the way that compressed JPEGs make up the pixels or the way that compressed video makes up chunks of image frames.

它是一种有用的技术,但它不是革命性的,也不是自1950年代以来最具影响力的东西。
View/Hide Original English

It's useful tech, but it's not revolutionary, nor is it the most impactful thing since the 1950s.

它只是我们多年来一直在做的事情的又一次迭代,远不如我今天谈到的许多变化重要。
View/Hide Original English

It's just another iteration of what we've been doing for years and not nearly as important as a lot of the changes I've talked about today.

现在,我应该说我不是第一个提出这个想法的人。
View/Hide Original English

Now, I should say I'm not the first to come up with this idea.

我第一次接触这个想法是在特德·姜(**Ted Chiang**: 华裔美国科幻小说作家)在《纽约客》(**The New Yorker**: 美国著名杂志)上发表的一篇题为《ChatGPT是网络的模糊JPEG》(**ChatGPT Is a Blurry JPEG of the Web**: 特德·姜关于ChatGPT本质的著名文章)的散文中。
View/Hide Original English

I first encountered the idea in an essay from Ted Ching called "ChatGPT ChatGPT is a blurry JPEG of the web," which was in the New Yorker.

链接在下方。
View/Hide Original English

Link below.

在我看来,特德是我的同代人中最优秀的科幻短篇小说作家,你可能看过或听说过2016年的电影《降临》(**Arrival**: 2016年科幻电影),它改编自他的短篇小说《你一生的故事》(**Story of Your Life**: 特德·姜的科幻短篇小说)。
View/Hide Original English

Ted is, in my opinion, by far the best speculative short story writer of my generation, you might have seen or heard of the movie Arrival from 2016, which is based on his short story "story of your life."

如果你还没有读过他的作品,我强烈推荐你去读一读。
View/Hide Original English

If you haven't read his stuff, I highly recommend you give him a read.

他简直太棒了。
View/Hide Original English

He's just amazing.

LLM作为有损压缩的含义与局限

LLM作为有损压缩的含义很有趣。

View/Hide Original English

So the implications of LLMs being lossy compression are interesting.

首先,你拥有整个互联网的可搜索图像,并且它能装进硬盘或大型GPU的RAM中,这绝对令人震惊。
View/Hide Original English

First off, the idea that you have a searchable image of the whole internet that will fit on a hard drive or in the RAM of a large GPU is absolutely mind-blowing.

现在,它绝不是一个忠实的图像。
View/Hide Original English

Now, it's not a faithful image by any stretch.

事实上,它是一个相当糟糕的副本。
View/Hide Original English

And in fact, it's a pretty crappy copy.

但根据你想要做什么,这通常是没问题的。
View/Hide Original English

But depending on what you're trying to do, that's often fine.

如果你想写小说,那太棒了。
View/Hide Original English

If you want to write fiction, that's great.

如果你想有一个想象中的笔友,它也很适合。
View/Hide Original English

If you want to have an imaginary pen pal, it's a good fit for that.

如果你想要商业提案的**样板文本**(Boilerplate text: 可重复使用的标准文本)或在**Stack Overflow**(Stack Overflow: 程序员问答社区)上有35个重复条目的**样板代码**(Boilerplate code: 可重复使用的标准代码),它是一种快速简便的获取方式。
View/Hide Original English

If you want boilerplate text for a business proposal or boilerplate code that has 35 duplicate entries on Stack Overflow, it's a fast and easy way to get it.

如果你想要解决技术面试编程谜题或国际数学竞赛问题或大学或法律入学考试或专业认证考试的学习指南摘录,它都能满足你。
View/Hide Original English

If you want an excerpt from a study guide for solving a tech interview programming riddle or an international mathematics competition problem or college or law entrance exams or professional certifications exams, it's got you covered.

此外,由于信息编码的方式,它在处理模糊、不明确或非结构化查询方面做得非常出色,这降低了人们使用这项技术所需的技能门槛。
View/Hide Original English

Also, because of the way the information is encoded, it does an amazing job of handling vague and vigorous or unstructured queries, which lowers the bar for what skills a person needs to have to put this tech to you.

另一方面,如果错误的后果很高,无论是医疗诊断、心理健康治疗、法律法庭文件中的判例法、刑事案件中的证据、需要防范黑客攻击的代码或杀手机器人的指令,它都不是一个好的选择。
View/Hide Original English

On the other hand, if the consequences for mistakes are high, be it for medical diagnoses, mental health therapy, case law in legal court documents, evidence and criminal cases, code that needs to be secure against hackers or commands for killer robots, it's not a good choice at all.

我们以前也经历过这种情况。
View/Hide Original English

We've been down this road before.

当**X射线**(X-rays: 一种电磁辐射)和**CT扫描**(CAT scans: 计算机断层扫描,一种医学影像技术)数字化时,人们非常担心如果图像使用**有损压缩**存储,可能会导致误诊。
View/Hide Original English

When x-rays and CAT scans went digital, there was a lot of concern about potential misdiagnosis if the images were stored using lossy compression.

自那时以来,已经有大量文献和研究探讨了这个问题。
View/Hide Original English

And there's been a lot of literature and studies examining that problem since.

遗憾的是,这个行业一直在假装,而且显然有时甚至相信,这些东西是活生生的生物或诸如此类的废话。
View/Hide Original English

The shame is the industry keeps pretending, and apparently sometimes even believing, that these things are living beings or some such crap.

这使我们无法就这项技术适合或不适合哪些用途进行诚实的对话。
View/Hide Original English

It keeps us from being able to have honest conversations about what uses this technology are good fits for or poor fits for.

关于**LLM**以及它们如何只是带有非结构化查询语言的**有损数据库搜索器**,我还有很多话要说。
View/Hide Original English

There's a whole lot more I have to say about LLMs and how they are just lossy database searchers with unstructured query language.

但这需要我深入探讨**数据仓库**、**数据湖**(Data Lakes: 存储原始格式的大量数据的存储库),并谈论我参与过的一些大型数据项目。
View/Hide Original English

But it requires me to get in the weeds about data warehouses, data lakes, and talk about some of the large data projects I've worked on.

所以我将把它放在我的副频道的一个视频中。
View/Hide Original English

So I'm putting that in a video on my secondary channel.

在这个视频发布时,我还没有完成那个视频,但它将是我上传到那个频道的下一个视频。
View/Hide Original English

I'm not done with that video at this time that this publishes, but it will be on the next one I upload on that channel.

一旦我上传了它,我会在下面的描述中放一个链接。
View/Hide Original English

And once I get it uploaded, I'll put a link to it in the description below.

如果你想知道那个视频或那里上传的其他新视频何时发布,请随时订阅我的另一个频道。
View/Hide Original English

Feel free to subscribe to my other channel if you want to know when that video or other new videos go up there.

这个视频的描述中有一个指向另一个频道的链接。
View/Hide Original English

There's a link to the other channel in this video's description.

顺便说一句,如果你还没有订阅这个频道,也请随时订阅。
View/Hide Original English

And while you're at it, feel free to subscribe to this channel too.

历史上的就业颠覆与AI的营销泡沫

我意识到对于很多人来说,当前的AI浪潮即使没有别的,也从就业角度来看是独一无二且极具颠覆性的。

View/Hide Original English

I realized for a lot of people that the current crop of AI is seen unique and incredibly disruptive from an employment standpoint of nothing else.

但从历史的角度来看,这根本不是前所未有的。
View/Hide Original English

But from the point of view of history, this isn't unprecedented at all.

你能想象有多少文书工作和时间花在文件柜和手动簿记上,而这些都被**SQL数据库**取代了吗?
View/Hide Original English

Can you imagine how much paperwork and time was spent with filing cabinets and manual bookkeeping that was all replaced by SQL databases?

我见过高达美国劳动力总数10%的估计。
View/Hide Original English

I've seen estimates as high as 10% of the entire US labor force.

但不幸的是,我找不到一个好的来源来证实这一点。
View/Hide Original English

But unfortunately, I can't find a good source for that.

职位名称和分类太多,而且太模糊,无法确定。
View/Hide Original English

There are just too many job titles and classifications and it's too vague to know for sure.

但这里有一个我有可靠来源的数字,供你思考。
View/Hide Original English

But here's a number I do have a good source for: just so you have something to think about.

猜猜曾经有多少人仅仅受雇为电话接线员?
View/Hide Original English

Guess how many people used to be employed just as telephone operators?

大约140万人,其中约40万人为电话公司本身工作,约100万人为其他雇主工作,而当时美国劳动力总数约为5900万人。
View/Hide Original English

About 1.4 million people, about 400,000 of those working for the phone coming itself and about a million working for other employees out of a total US labor force of roughly 59 million people.

在高峰期,美国每42个工作人口中就有一个是电话接线员。
View/Hide Original English

That's one of every 42 working people in the US or telephone operators at the peak.

这只是一个现在已经完全自动化消失的职业,而许多职业都因技术进步而大幅减少。
View/Hide Original English

And that's just one single occupation that's now been completely automated out of existence out of the many occupations that were radically reduced by the march of technology.

当前**生成式AI**浪潮与之前几次技术采纳周期的最大区别不在于颠覆程度。
View/Hide Original English

The biggest difference between the generative AI wave now and the several previous technology adoption cycles isn't the level of disruption.

而是营销、宣传和**煤气灯效应**(Gaslighting: 一种心理操纵形式,让受害者怀疑自己的记忆、感知或理智)的程度。
View/Hide Original English

It's the levels of marketing, propaganda and gaslighting.

哦,还有炒作。
View/Hide Original English

Oh, and the hype.

别忘了炒作。
View/Hide Original English

Don't forget the hype.

嗯,好像你真的能忘记炒作一样。
View/Hide Original English

Well, as if you could forget the hype.

所以从现在开始,当你谈论或思考**生成式AI**的好坏应用时,试着把它想象成互联网的一个粗糙、模糊的副本,而不是某种神话般的科幻小说中即将出现的超人,看看这是否能帮助你更好地理解它。
View/Hide Original English

So from now on, when you're talking about or thinking about good or bad applications of generative AI, try thinking about it as a crappy, blurry copy of the internet instead of some mythical sci-fi soon to be superhuman and see if that doesn't help you put it in a better context.

感谢观看。
View/Hide Original English

Thanks for watching.

大家在外要小心。
View/Hide Original English

Let's be careful out there.