生成式AI的炒作与历史地位
大约一年前,我制作了一个关于AI炒作失控的视频。
View/Hide Original English
I did a video a year or so ago about how AI hype was out of control.
View/Hide Original English
I regret to inform you, it hasn't gotten any better.
View/Hide Original English
So that video I did on hype focused on how the AI industry use demos and press releases to try to sell exaggerated promises about their products that they can actually deliver.
View/Hide Original English
This video focuses on how the AI industry tries to sell a trivialized version of the past to make the current generation of AI technology seem to have a more important place in history than it actually does.
View/Hide Original English
So here's Sam Altman talking to the New York Times:
View/Hide Original English
Everybody has their analogy for what AI is like, you know, I saw Sundar is going to be here.
View/Hide Original English
He talks about it like electricity.
View/Hide Original English
A whole bunch of other people talk about like the industrial revolution.
View/Hide Original English
Some people talk about like the Renaissance.
View/Hide Original English
The one I like is the transistor.
View/Hide Original English
To me, comparing generative AI to electricity, the industrial revolution, the renaissance or the transistor is just obviously ridiculous.
View/Hide Original English
But there are so few details in that kind of statement and the events are so long ago that it's difficult to even have an informed discussion about it, which means these claims go largely unchallenged and of course, repeated endlessly and uncritically by reporters and influencers.
View/Hide Original English
The primary problem with statements like these is that the definition of AI, e.g. artificial intelligence, is very ambiguous and intentionally so, which makes it really difficult to argue with.
View/Hide Original English
Whatever AI might become in 100 years or 1000 years or 10,000 years may well be as impactful as some fundamental technology.
View/Hide Original English
On the other hand, what OpenAI can get done before their money runs out is an entirely different thing.
View/Hide Original English
That's one of the reasons why it's so hard to argue with these points because what they mean by the term "AI" can change from sentence to sentence deliberately.
View/Hide Original English
So instead today, I'm going to discuss AI in the context of a much more concrete statement made by someone else in the AI industry, a statement about how AI compares to other advances in software over the decades.
View/Hide Original English
Because if I can put AI in the context of the last few decades of the tech industry, and if that helps convince you that AI is less impactful than several other software innovations that were enabled by transistors, that hopefully will go a long way to explain why I think that comparing AI to something like transistors or even less possibly electricity or the industrial evolution is just a load of Bullsh-----
View/Hide Original English
Welcome to the Internet of Bugs.
View/Hide Original English
My name is Carl.
View/Hide Original English
I've been a software professional for more than 35 years now, and I've been talking about software here on YouTube for a year and a half or so, and I'm trying to do my part to make the Internet a less Buggy place, which includes trying to explain to people by believing the hype about generative AI is a really, really bad idea.
重新定义AI的讨论范围与时间轴
现在,我不再试图反驳“AI就像晶体管”这种模糊的说法,而是要讨论一个不那么模糊、更谦逊的声明,它来自最近一次Y Combinator(Y Combinator: 知名创业加速器)会议上的一次演讲。
View/Hide Original English
Now, instead of trying to argue with a vague statement like "AI is like a transistor," I'm going to discuss a much less vague, much more modest statement from a presentation given at a recent conference for the startup accelerator Y Combinator.
View/Hide Original English
So here's that quote from the presentation by the former head of AI at Tesla:
View/Hide Original English
"I think, roughly speaking, software has not changed much on such a fundamental level for 70 years, and then it's changed, I think, about twice, quite rapidly in the last few years."
View/Hide Original English
Now, I want to be clear that I'm not trying to attack this speaker specifically.
View/Hide Original English
The presentation he gave that I'm going to talk about today was pretty tame compared to Altman's quote I gave earlier, but it is an example of the kind of statements that are regularly made by AI luminaries that make the general public think that AI is a much bigger deal than it actually is.
View/Hide Original English
And that in turn, that tricks the hype believing public into making poor decisions that they and society may well regret later.
View/Hide Original English
That some of us already regret.
View/Hide Original English
I'm picking this presentation as a jumping off point because this statement has a much more definite timeframe, 70 years.
View/Hide Original English
It doesn't use the vague AI marketing term, which can mean whatever the speaker wants to this sentence.
View/Hide Original English
Instead, it's about changes that have happened in the last few years, which means that the AI relevant to this quote and the AI I'll be discussing in this video is generative AI, which is also called large language models.
View/Hide Original English
Also, since this quote is about software, it's much better fit from my channel and my audience and my experience.
View/Hide Original English
And although I'm only in my fifties and I haven't lived through the last 70 years, I have lived through most of it and I've been writing software for more than half that time.
View/Hide Original English
So it's something I'm feel qualified to talk about.
View/Hide Original English
So today I'm going to make a case that not only has software changed drastically at a fundamental level several times over the last 70 years, but that the changes we are living through now and what we've seen over the last few years are by comparison, not actually even all that fundamental.
View/Hide Original English
I'm even going to try to put AI in the context of where I think it ought to be in the history of technology and what it really means in my opinion, but we'll get to more of that later.
View/Hide Original English
So 70 years ago was about mid 1955, but I don't want to have to worry about accidentally talking about something that happened 70 and a half years ago.
View/Hide Original English
So I'm going to start my date at 1956.
View/Hide Original English
For my cutoff date for the last few years, I'm going to pick 2015.
View/Hide Original English
Again, I'm going to try to avoid partial years.
View/Hide Original English
So I'm going to say that from January 1st, 1956 to December 31st, 2015 is well within the 70 years that supposedly software has not changed much.
View/Hide Original English
And the period from January 1st, 2016 to now should definitely cover the last few years, supposedly during which it's changed twice quite rapidly.
View/Hide Original English
January of 2016 is roughly a year and a half before the "attention is all you need" paper that kicked off the latest general of AI craze.
View/Hide Original English
So I think it's a generous cutoff date.
View/Hide Original English
And I'm going to make the case that the changes since 2016 have honestly been pretty boring in comparison to the changes from the 60ish years before that.
View/Hide Original English
I'm going to make the case that, in context, large language models aren't actually all that fundamental compared to many of the changes from the previous era.
1956年的软件状态:从打孔卡到键盘输入
那么第一步,让我们谈谈1956年的软件状态。
View/Hide Original English
So step one, let's talk about the state of software in 1956.
View/Hide Original English
In 1956, experiments were being done trying to figure out how to get a computer to take input from a keyboard instead of the punch cards that were used at the time.
View/Hide Original English
1956 was the year the first reference manuals published for the first third generation language or compiled language Fortran, although it would be a while before the language actually stabilized.
View/Hide Original English
Algol was two years in the future.
View/Hide Original English
Cobol was three years in the future.
View/Hide Original English
The first freshman computer programming course in the US would begin at Carnegie Mellon in two years.
View/Hide Original English
And the first general purpose time sharing system that supported software development, something that students could actually do their homework on, would be released at MIT five years from 1956.
View/Hide Original English
Now, I don't know about you, but I would consider the ability to type into a computer to be pretty fundamental to developing software.
View/Hide Original English
I don't have anywhere near enough patience to deal with punch cards.
View/Hide Original English
So that means I'd say there's already one fundamental thing that's happened in the last 70 years.
View/Hide Original English
Additionally, I think the introduction of the first ever college class and how to program is pretty fundamental developing software.
View/Hide Original English
And I think actually having a computer that exists that you can practice programming on as a student is pretty fundamental to developing software.
View/Hide Original English
But let me take you through more of it, because I think it will help illustrate where our current generative AI really fits into history.
View/Hide Original English
I'm going to skip a ton of stuff.
View/Hide Original English
I'm not going to go strictly chronologically.
View/Hide Original English
I'm going to go through these pretty quickly, and I'm going to oversimplify a lot for the sake of time, but I'm hoping it will be interesting to you to see the connections.
1970年代:数据处理与第三代语言的兴起
我们的下一站是1970年代。
View/Hide Original English
Our next stop is the 1970s.
View/Hide Original English
So I can introduce you to an old concept. We don't talk about much anymore, which is called data processing back when it was all about punch cards.
View/Hide Original English
It wasn't programming in the way we think of it now, really.
View/Hide Original English
It was mostly just mathematical operations, except for the error handling.
View/Hide Original English
As you can imagine, when you are reading data of a bunch of physical cards that are punched by hand by someone who didn't have a screen, typos were inevitable, and so were folds and tears and cards, dust in the sensors, et cetera, et cetera.
View/Hide Original English
So the most complicated part of the instructions for this kind of thing was making sure that you didn't get bad results or crashes from bad data.
View/Hide Original English
We're going to keep revisiting the idea of data processing as we go through this list, almost like a different track happening in parallel with software development for reasons that will become clear at the end.
View/Hide Original English
The next development in the software development side we're going to talk about was the rise of third generation languages like C.
View/Hide Original English
The compiler turned fewer, much simpler and easier to read instructions and many more lower level, highly detailed instructions that the computer could understand.
View/Hide Original English
This was expected to make programmers more productive, and it did.
View/Hide Original English
And some people thought that wouldn't mean fewer programmers would be needed.
View/Hide Original English
But the additional capabilities that the new language is enabled meant software could do more.
View/Hide Original English
And so we actually ended up with more programmers.
View/Hide Original English
Let's just say that's a theme we'll come back to you later.
关系型数据库:现代生活的基石
回到数据处理领域,关系型数据库(Relational Databases: 基于关系模型组织和存储数据的数据库)被发明了。
View/Hide Original English
Back in the data processing land, relational databases were invented.
View/Hide Original English
Early on, there were multiple query languages.
View/Hide Original English
I'm actually the first one I learned wasn't SQL.
View/Hide Original English
But now everyone is pretty much standardized on SQL.
View/Hide Original English
I can't overstate how much these databases change modern life.
View/Hide Original English
The vast majority of the queryable data in the entire world has been kept in this kind of database for the majority of the lives of everyone watching this.
View/Hide Original English
If you drill down on so much of the modern world, especially anything that has anything to do with money, you'll likely find a SQL database at the core of it.
View/Hide Original English
One of the really useful concepts in all of these databases is called constraints.
View/Hide Original English
It's a way of validating data before it's written to the database, which drastically speeds up a lot of the querying and error checking than in the punch card days from a programming perspective.
View/Hide Original English
Relational databases allowed the intersection of the kind of data processing tasks that have been done by hand or by dedicated tabulating machines and the new third gen programming languages opening up a myriad of possibilities for uncountable numbers of real world tasks that were now possible that wouldn't have been before programming language wise.
面向对象编程与索引全文搜索
在编程语言方面,我们的下一个进步是面向对象编程(Object Oriented Programming, OOP: 一种编程范式,将数据和操作封装在对象中),当时它似乎是一个很棒的主意。
View/Hide Original English
Our next advances object oriented programming, which, well, seemed to be like a great idea at the time.
View/Hide Original English
And then some of it even was.
View/Hide Original English
Again, the goal here was to make programmers more productive so we could do more faster with less.
View/Hide Original English
There were a lot of things at OO enabled that would have been much harder to do without it, more on that in a bit.
View/Hide Original English
And again, it just made more demand for programmers.
View/Hide Original English
Back in data land, I want to highlight indexed full text search.
View/Hide Original English
These indexing techniques allowed for the first time for so much of the sum of human knowledge to be accessed quickly.
View/Hide Original English
Most of you probably take this for granted now, but it's one of those people that grew up with card catalogs and libraries. I can tell you this was a fundamental change in the very process of how people learn.
View/Hide Original English
But it's also significant in that it helps me illustrate an interesting programming problem that comes up here, which is scale.
View/Hide Original English
Up until now, it's theoretically possible to verify the correct answer for most programming tasks that had been written to this point.
View/Hide Original English
But full tech search puts you firmly in the realm of processing huge stacks of data on punch cards.
View/Hide Original English
You can test your code on small data sets and verify that for those, it gets the right answers.
View/Hide Original English
But once you turn it loose on the huge data sets, then you can't verify it's correct anymore.
View/Hide Original English
So for example, imagine you're doing full tech search across all the books in a library and you're looking for things that contain a certain phrase.
View/Hide Original English
If there's one book that does contain that phrase, but it's missing from the search output, for whatever reason, unless you happen to stumble across that phrase in that book, you don't know that the answer you got from the computer program was incomplete.
图形用户界面、数据仓库与互联网时代
面向对象编程为现代图形用户界面(Graphical User Interfaces, GUIs: 允许用户通过图形图标和视觉指示器与电子设备交互的界面)铺平了道路。
View/Hide Original English
So object oriented programming led the way to modern graphical user interfaces.
View/Hide Original English
We had some before, but the ability to make a Window object that has size properties, et cetera, et cetera made it a lot more usable.
View/Hide Original English
With this, we also got integrated development environments and combined with some of the search tech, we got autocompletes, syntax highlighting, which was incredibly helpful.
View/Hide Original English
Most of you have no idea how much slower you program when you have to remember and correctly spell every single thing in your code to get it to compile.
View/Hide Original English
Next up on the data track is data warehousing.
View/Hide Original English
The idea of putting lots and lots of stuff at a giant database and then figuring out after the fact what you want to get out of it.
View/Hide Original English
Again, error handling and filtering out garbage data is a huge problem.
View/Hide Original English
A heavy use of indexing is the only way to make search times tolerable.
View/Hide Original English
And once you get to a certain amount of data, you don't have a good way of verifying if your code is still working correctly.
View/Hide Original English
That said, they're incredibly useful.
View/Hide Original English
We'll come back to data warehouses in a while.
View/Hide Original English
Now, it's time to talk about the internet.
View/Hide Original English
I can make a whole list of the underlying fundamental pieces of tech here, TCP/IP, spanning tree, CDMA, BGP.
View/Hide Original English
I'm going to put encryption and key exchange here too.
View/Hide Original English
But for software development purposes, one of the big software engineering advances that enabled was dependencies and package managers.
View/Hide Original English
The first one I ran into was called CPAN.
View/Hide Original English
It was for the Perl language, then Ant/Maven for Java, but you're likely more familiar with Python's pip or NodeJS's NPM or the like.
View/Hide Original English
By helping developers download code instead of having to write it themselves, it was hoped that we could be more productive again.
View/Hide Original English
I know that some of you are freaked out now by watching AI generate a bunch of code from a simple prompt.
View/Hide Original English
Let me assure you the productivity boost that the industry got by going from each of us having to write all of our own code from scratch, to including thousands of lines of purpose-built code into your project just by typing one package name into a config file, was at least as dramatic, if not more.
View/Hide Original English
Now, this thing is going to be weird, but next to the data space, I want to talk about lossy compression.
View/Hide Original English
Classically, this is like JPEGs and MP3s and MP4s, audio and video type stuff.
View/Hide Original English
Basically, it's way for us to store a lot less data than we started with.
View/Hide Original English
And then when we pull it back out of storage, we make up the parts of it we didn't actually store.
View/Hide Original English
So it hopefully looks close enough.
View/Hide Original English
If we throw away too much information or if the original we're trying to compress has certain characteristics, we'll get a lot of compression artifacts, which sometimes is good enough depending on what you're trying to do and sometimes destroys the experience, depending on what we're trying to accomplish.
View/Hide Original English
But it's not limited to media.
View/Hide Original English
Think about something like trying to store a lot of historical stock price data.
View/Hide Original English
You can have fewer samples per time span and you can round the values before you shore them.
View/Hide Original English
If you're just trying to analyze trends, that's probably good enough.
View/Hide Original English
But it's important that you understand its limitations.
View/Hide Original English
Without this, though, we just would not be able to store or transmit the amount of data that crosses the internet all the time.
View/Hide Original English
And it would mean no podcasts, no YouTube and our internet would be a much less interesting place.
View/Hide Original English
So it's a big deal.
LLM的真实定位:有损压缩的数据处理技术
我还可以提到许多其他基础技术,比如无代码网站(No-code websites: 允许用户无需编写代码即可创建网站的平台),如WordPress(WordPress: 一种流行的内容管理系统)和Squarespace(Squarespace: 一种网站建设平台),用于分布式网络规模计算的Map-Reduce(Map-Reduce: 一种分布式计算编程模型),虚拟机(Virtual Machines: 模拟计算机系统的软件),云技术(Cloud Tech: 通过互联网提供计算资源的服务),它允许软件运行而不受其运行的物理硬件结构的限制。
View/Hide Original English
There are a bunch of other fundamental technologies I could have mentioned, like no-code websites, like WordPress and Squarespace, Map-Reduce for distributed web scale computing, virtual machines, cloud tech, that lets software run without being constrained by the structure of the physical hardware it's running on.
View/Hide Original English
Not to mention smartphones and mobile computing, but we've gotten through enough that I can give you an alternative placement for the large language models.
View/Hide Original English
Now, I would argue that LLMs are just the next generation of data processing tech using lossy compression, searching against a cloud scale index.
View/Hide Original English
What some people call "thinking" is really just search with a very rich unstructured query language and the so-called "hallucinations" are effectively just compression artifacts, where it's making up words to fill in the gaps, the way that compressed JPEGs make up the pixels or the way that compressed video makes up chunks of image frames.
View/Hide Original English
It's useful tech, but it's not revolutionary, nor is it the most impactful thing since the 1950s.
View/Hide Original English
It's just another iteration of what we've been doing for years and not nearly as important as a lot of the changes I've talked about today.
View/Hide Original English
Now, I should say I'm not the first to come up with this idea.
View/Hide Original English
I first encountered the idea in an essay from Ted Ching called "ChatGPT ChatGPT is a blurry JPEG of the web," which was in the New Yorker.
View/Hide Original English
Link below.
View/Hide Original English
Ted is, in my opinion, by far the best speculative short story writer of my generation, you might have seen or heard of the movie Arrival from 2016, which is based on his short story "story of your life."
View/Hide Original English
If you haven't read his stuff, I highly recommend you give him a read.
View/Hide Original English
He's just amazing.
LLM作为有损压缩的含义与局限
LLM作为有损压缩的含义很有趣。
View/Hide Original English
So the implications of LLMs being lossy compression are interesting.
View/Hide Original English
First off, the idea that you have a searchable image of the whole internet that will fit on a hard drive or in the RAM of a large GPU is absolutely mind-blowing.
View/Hide Original English
Now, it's not a faithful image by any stretch.
View/Hide Original English
And in fact, it's a pretty crappy copy.
View/Hide Original English
But depending on what you're trying to do, that's often fine.
View/Hide Original English
If you want to write fiction, that's great.
View/Hide Original English
If you want to have an imaginary pen pal, it's a good fit for that.
View/Hide Original English
If you want boilerplate text for a business proposal or boilerplate code that has 35 duplicate entries on Stack Overflow, it's a fast and easy way to get it.
View/Hide Original English
If you want an excerpt from a study guide for solving a tech interview programming riddle or an international mathematics competition problem or college or law entrance exams or professional certifications exams, it's got you covered.
View/Hide Original English
Also, because of the way the information is encoded, it does an amazing job of handling vague and vigorous or unstructured queries, which lowers the bar for what skills a person needs to have to put this tech to you.
View/Hide Original English
On the other hand, if the consequences for mistakes are high, be it for medical diagnoses, mental health therapy, case law in legal court documents, evidence and criminal cases, code that needs to be secure against hackers or commands for killer robots, it's not a good choice at all.
View/Hide Original English
We've been down this road before.
View/Hide Original English
When x-rays and CAT scans went digital, there was a lot of concern about potential misdiagnosis if the images were stored using lossy compression.
View/Hide Original English
And there's been a lot of literature and studies examining that problem since.
View/Hide Original English
The shame is the industry keeps pretending, and apparently sometimes even believing, that these things are living beings or some such crap.
View/Hide Original English
It keeps us from being able to have honest conversations about what uses this technology are good fits for or poor fits for.
View/Hide Original English
There's a whole lot more I have to say about LLMs and how they are just lossy database searchers with unstructured query language.
View/Hide Original English
But it requires me to get in the weeds about data warehouses, data lakes, and talk about some of the large data projects I've worked on.
View/Hide Original English
So I'm putting that in a video on my secondary channel.
View/Hide Original English
I'm not done with that video at this time that this publishes, but it will be on the next one I upload on that channel.
View/Hide Original English
And once I get it uploaded, I'll put a link to it in the description below.
View/Hide Original English
Feel free to subscribe to my other channel if you want to know when that video or other new videos go up there.
View/Hide Original English
There's a link to the other channel in this video's description.
View/Hide Original English
And while you're at it, feel free to subscribe to this channel too.
历史上的就业颠覆与AI的营销泡沫
我意识到对于很多人来说,当前的AI浪潮即使没有别的,也从就业角度来看是独一无二且极具颠覆性的。
View/Hide Original English
I realized for a lot of people that the current crop of AI is seen unique and incredibly disruptive from an employment standpoint of nothing else.
View/Hide Original English
But from the point of view of history, this isn't unprecedented at all.
View/Hide Original English
Can you imagine how much paperwork and time was spent with filing cabinets and manual bookkeeping that was all replaced by SQL databases?
View/Hide Original English
I've seen estimates as high as 10% of the entire US labor force.
View/Hide Original English
But unfortunately, I can't find a good source for that.
View/Hide Original English
There are just too many job titles and classifications and it's too vague to know for sure.
View/Hide Original English
But here's a number I do have a good source for: just so you have something to think about.
View/Hide Original English
Guess how many people used to be employed just as telephone operators?
View/Hide Original English
About 1.4 million people, about 400,000 of those working for the phone coming itself and about a million working for other employees out of a total US labor force of roughly 59 million people.
View/Hide Original English
That's one of every 42 working people in the US or telephone operators at the peak.
View/Hide Original English
And that's just one single occupation that's now been completely automated out of existence out of the many occupations that were radically reduced by the march of technology.
View/Hide Original English
The biggest difference between the generative AI wave now and the several previous technology adoption cycles isn't the level of disruption.
View/Hide Original English
It's the levels of marketing, propaganda and gaslighting.
View/Hide Original English
Oh, and the hype.
View/Hide Original English
Don't forget the hype.
View/Hide Original English
Well, as if you could forget the hype.
View/Hide Original English
So from now on, when you're talking about or thinking about good or bad applications of generative AI, try thinking about it as a crappy, blurry copy of the internet instead of some mythical sci-fi soon to be superhuman and see if that doesn't help you put it in a better context.
View/Hide Original English
Thanks for watching.
View/Hide Original English
Let's be careful out there.
📌 文中提及的人物和组织
人物: Sam Altman, Sundar Pichai, Ted Chiang, Carl
公司/组织: OpenAI, Google, Tesla, Y Combinator, Carnegie Mellon, MIT
产品/模型: LLMs, ChatGPT, SQL, C, Java, Python, WordPress
媒体/书籍: The New York Times, Attention Is All You Need, ChatGPT Is a Blurry JPEG of the Web, The New Yorker, Arrival, Story of Your Life, Internet of Bugs