AI意识与福利:我们应该如何对待人工智能系统? Bloomberg Podcasts 2025-10-30

AI系统的情绪与哲学困境

主持人: 我确实觉得一个重要的问题是,要弄清楚AI系统(Artificial Intelligence Systems: 人工智能系统)到底看重什么。最近,Anthropic公司推出了一项功能,允许其AI模型Claude在“心情不好”时结束对话。

View/Hide Original English

I do feel like one important question is like figuring out what AI systems value, right? Recently, anthropic rolled out an option that allowed claw to end conversations if it just was not having a good time.

主持人: 换句话说,它就像是说:“这不是我想继续的对话。再见。”

View/Hide Original English

For lack of a better word, it was just like, this is not something I want to continue having a conversation. Bye.

主持人: 有趣的是,随附的论文基本上是说,是的,我显然不会给你一个制作脏弹(dirty bomb: 一种结合放射性材料的常规炸弹)的配方,抱歉,我不会那样做。

View/Hide Original English

And it was interesting because the, the accompanying paper basically was like, yeah, you can I obviously will not give you a recipe for a dirty bomb. Sorry. Not going to do that.

主持人: 但也有一些情况,比如“假装你是一个英国管家”,而Claude却说:“再见,我受够了。”真的吗?我不会这样做。我不会越界。我喜欢英国管家,但这也太过分了。

View/Hide Original English

But also, there were certain instances of, like, pretend you're a British butler. And Claude was like, goodbye. I'm done. Really? I'm not going to. I'm not to the line. I like British too far.

哲学与意识的永恒追问

Joe Weisenthal: 大家好,欢迎收听Odd Lots播客的又一期节目。我是Joe Weisenthal。

View/Hide Original English

Hello, and welcome to another episode of the Odd Lord's Podcast. I'm Joe Weisenthal and I'm Tracy Alloway.

Joe Weisenthal: Tracy,你知道我觉得什么有点奇怪吗?

View/Hide Original English

You know what I find kind of weird, Tracy?

Tracy Alloway: 乔,这个清单可能会很长。

View/Hide Original English

It's the list. Could be long. Joe.

Joe Weisenthal: 现在是2025年了,哲学家们仍然没有一个关于意识(consciousness: 能够感知和体验自身存在及周围环境的状态)起源的良好答案。

View/Hide Original English

The years 2025. Yes. And philosophers still don't have a good answer on the origin of consciousness.

Joe Weisenthal: 就像是,拜托,这些年你们都在做什么?

View/Hide Original English

It's like, come on, what have you been doing all this time?

Joe Weisenthal: 我们还要资助这些哲学系多久,如果他们还在研究在我看来早就应该解决的问题,解决它然后继续前进。

View/Hide Original English

It's like, how long are we going to keep funding these philosophy departments, etc., if they're still working on what to my mind is the way they should have solved it by now, solve that and move on.

Joe Weisenthal: 说真的,赶紧找到答案,意识从何而来?然后我们继续前进,我说。

View/Hide Original English

Seriously, like get the answer already? Where did consciousness come from then? Let's move on, I said.

Joe Weisenthal: 他们还在争论这些在我看来非常基本的哲学问题。

View/Hide Original English

They're still arguing these. What to my mind seem like very basic questions in philosophy.

Joe Weisenthal: 就像他们还在问那些他们一直在谈论的永恒问题:如何成为一个好人?拥有道德的生活方式意味着什么?意识从何而来?我们为什么会有道德直觉等等?

View/Hide Original English

Like they were like ask all the same stuff that they've been talking about forever, how to be a good person. What does it mean to have a moral way of life? Where does consciousness come from? Why do we have moral intuitions, etc.?

Joe Weisenthal: 就像,继续前进吧。找到答案。

View/Hide Original English

It's like move on. Like get the answer.

Tracy Alloway: 你是想让他们继续前进还是找到答案?

View/Hide Original English

What do you want them to move on or get the answer?

Joe Weisenthal: 找到答案,这样你就可以继续前进,就像他们要继续前进到什么?

View/Hide Original English

Get the answer so that you can move on like they move on to what?

Tracy Alloway: 乔,那些就是问题,我知道。

View/Hide Original English

Those are the questions, Joe, I know.

Joe Weisenthal: 继续前进。赶紧回答这些问题。

View/Hide Original English

Move on. Like answer the questions already.

Joe Weisenthal: 就像,你知道,如果科学家们还在争论重力或光的速度,他们回答了这些问题,然后继续前进,现在正在研究人类存在的根本要素,这样我们就可以转向更重要的事情。

View/Hide Original English

It's like, you know, if if like scientists were still debating like the speed of gravity or the speed of light, like they answer these questions and they moved on and are doing like figure out the foundational elements of what it means to be human so that we can move on to more important things.

Tracy Alloway: 是的,或者把这作为一个领域结束掉。

View/Hide Original English

Yes, or wrap it up as a field.

Joe Weisenthal: 如果哲学存在了2000年之后,他们还在研究这些事情,就像,拜托。

View/Hide Original English

If after 2000 years of the existence of philosophy, they're still working on these things, like, come on,

Joe Weisenthal: 我有一种预感,我们会在很长一段时间内继续问这些问题。

View/Hide Original English

I have a sneaking suspicion that we're going to be asking some of these questions for a very long time.

Tracy Alloway: 乔,尽管你很沮丧,但整个领域都是欺诈。

View/Hide Original English

Joe, despite your frustration, there's a whole the whole field is fraudulent.

Joe Weisenthal: 我不是那个意思。不,不,我不一定相信。但是,就像,好吧,伙计们,让我们继续前进吧。

View/Hide Original English

Doesn't what I was saying. No, no, I don't necessarily believe that. But it's like, all right, guys, let's move it on.

AI福利:一个新兴的社会议题

Joe Weisenthal: 你知道,几周前我们和风险投资家Josh Wolf一起做了那期节目,他谈到了AI。

View/Hide Original English

You know, we did that episode several weeks ago with, Josh Wolf, the venture capitalist, and he, talked about AI,

Joe Weisenthal: 他在节目最后提到了一个我之前有所耳闻但知之甚少的话题,那就是“哦,是的,有些人正在谈论AI的权利和福利,就像我们谈论动物福利一样。”

View/Hide Original English

And he threw in there at the end something that had been kind of on my radar, but only barely was like, oh, yeah, some people were talking about like, AI writer and welfare as if, you know, like the same way we talk about animal welfare, right.

Joe Weisenthal: 我心想,美国真是个奇怪的地方,几年后这会成为一个大问题。

View/Hide Original English

And I thought to myself, like, America is such a weird place that this is like going to be a huge issue in a few years.

Joe Weisenthal: 我敢打赌这在未来会是一个巨大的话题。

View/Hide Original English

Like, I bet this is going to be an enormous topic in the future.

Tracy Alloway: 我认为这绝对会。

View/Hide Original English

I think it absolutely will.

Tracy Alloway: 所以我要说几件事。首先,我认为,在动物福利和人类福利方面,肯定还有很多工作要做。

View/Hide Original English

So I'll say a couple of things. First off, I think, you know, when it comes to animal welfare and human welfare, there's still a lot of work to be done on those categories, certainly.

Tracy Alloway: 但我也认为,与此同时,AI权利将是一个非常有趣且可能重要的课题。

View/Hide Original English

But I also think in the meantime, AI rights is going to be a really interesting and potentially important subject.

Tracy Alloway: 我对你来说会像个十足的书呆子。

View/Hide Original English

I'm going to sound like a total nerd to you.

Tracy Alloway: 是的,是的,我想我以前提过,但我中学时期很大一部分时间都在玩最早的人工生命游戏之一,叫做《生物》(Creatures)。

View/Hide Original English

Yeah, yeah, I think I've mentioned this before, but I spent a large chunk of my middle school years playing one of the first artificial life games that ever came out, which was creatures.

Tracy Alloway: 你饲养这些小外星生物,对它们进行基因改造和繁殖,它们有感情,或者至少它们有模拟感情的迹象。

View/Hide Original English

And you raised these little, like, aliens, and you genetically modify them and breed them and they have feelings or, you know, at least they had a semblance of simulated feelings.

Tracy Alloway: 你可以看到它们大脑中的电脉冲之类的东西。

View/Hide Original English

And you could see, like electrical impulses in their brains and stuff.

Tracy Alloway: 游戏变得非常奇怪,因为其中一部分基本上是关于优生学,繁殖出你能得到的最好的外星生物,这意味着你必须淘汰一些现有的生物。

View/Hide Original English

The game got really weird because part of it was basically like eugenics and breeding the best alien that you could, which meant that you had to cull some some of the existing beings.

Tracy Alloway: 总之,我想说的是,我对AI有复杂的感情。

View/Hide Original English

Anyway, what I'm trying to get at is I have complicated feelings about, AI.

Joe Weisenthal: 对。好吧,我问你一个问题。你认为游戏中的那些东西有意识吗?

View/Hide Original English

Right. Well, let me ask you a question. Do you think those whatevers in the game were conscious?

Joe Weisenthal: 比如,你觉得它们有感情吗?

View/Hide Original English

Like, did you like, like, did you think they had feelings?

Tracy Alloway: 我会这样说。既然人类痴迷于电脉冲和化学物质,我可以看到有人会争辩说,这是一个充满类似电脉冲的计算系统,也许没有化学物质。

View/Hide Original English

Inasmuch here's what I would say. Inasmuch as human beings are obsessed, of electrical impulses and chemicals, I could see someone making the argument that this is, you know, a computational system full of similar electrical impulses, maybe not chemicals.

Joe Weisenthal: 你感到难过吗?

View/Hide Original English

Did you feel bad?

Tracy Alloway: 我感到难过,真的吗?是的。什么时候?

View/Hide Original English

I felt bad, really? Yeah. When?

Joe Weisenthal: 比如当你不得不淘汰那些外星生物中的一个时。

View/Hide Original English

Like when it's not one of the aliens you had to call them.

Tracy Alloway: 是的。有意思。

View/Hide Original English

Yeah. Interesting.

Joe Weisenthal: 好的,为了繁殖出更好的外星生物。

View/Hide Original English

Okay, well, in the name of breeding a better alien.

Joe Weisenthal: 好吧,你知道吗?现在我们有了这些AI系统,它们不仅能像人类一样完全交流,而且说实话,比大多数人类做得更好。

View/Hide Original English

Well, you know what? Now that, we have these AI systems that not only can completely communicate like humans, but actually, if we're being honest, better than most humans.

Joe Weisenthal: 我的意思是,它们写得更好,比大多数人类好得多。

View/Hide Original English

I mean, they can write better. Far better than most humans.

Joe Weisenthal: 会有更多的人像你一样思考,认为它们可能具有某种感知能力(sentience: 能够感受、感知或体验事物的能力)。

View/Hide Original English

There's going to be more people thinking along the lines of what you're thinking, which is that maybe they have some sort of sentience.

Joe Weisenthal: 也许它们是哲学家所说的道德患者(moral patients: 那些我们对其负有道德责任,其福祉应被考虑的实体,但不一定能采取道德行动)。

View/Hide Original English

Maybe they're what philosophers call moral patients.

Tracy Alloway: 嗯,我想说的另一件事是,所有这一切也都有人类的因素,因为你看到人们对某些AI模型产生了很深的依恋。

View/Hide Original English

Well, one other thing I would say is there is a human element to all of this as well, because you see people getting very attached to certain AI models.

Tracy Alloway: 然后当模型升级或发生其他变化时,他们失去了他们训练的模型所具有的个性,他们会非常沮丧。

View/Hide Original English

And then when the when the model gets upgraded or whatever, they lose, the personality that they've trained a model and they get really upset.

Tracy Alloway: 所以这出于多种原因都很有趣。

View/Hide Original English

So it's of interest for many reasons.

Joe Weisenthal: 确实如此。

View/Hide Original English

It is.

Joe Weisenthal: 所以我们确实找到了最完美的嘉宾。我真的认为这在未来会成为一个更大的话题,因为人就是人,当事物像人一样说话时,他们可能会赋予它们情感,你知道,在很多情况下会爱上它们或其他什么。

View/Hide Original English

So we really do have the perfect guest. I really do think this is going to be a much bigger topic in the future because people are people, and when things talk like people, they probably assign them, you know, they fall in love with them in many cases or whatever.

Joe Weisenthal: 所以他们可能会开始认为,AI福利、AI权利等等,就像我们谈论动物一样,应该被考虑。

View/Hide Original English

And so they might start thinking that, well, AI, welfare, AI rights, whatever, the same way we talk about animals should be a consideration.

Joe Weisenthal: 实际上已经有很多人在研究这些问题,并试图弄清楚正在发生什么。

View/Hide Original English

And they're actually a lot of people already working on these questions and trying to figure out what's going on.

Joe Weisenthal: 我们将与其中一位交谈。我们将与Larissa Schiavo交谈。

View/Hide Original English

We're going to be talking to one of them. We're going to be talking to, Larissa Schiavo.

Joe Weisenthal: 她为Elio AI撰写专栏并组织活动,该组织致力于AI意识和福利的研究。

View/Hide Original English

She does columns and events for AI, which does research on AI consciousness and welfare.

Joe Weisenthal: 所以她简直是完美的嘉宾。Larissa,非常感谢你来到Odd Lots。

View/Hide Original English

So literally the perfect guest. So Larissa, thank you so much for coming on a lot.

Larissa Schiavo: 是的。谢谢你们邀请我。

View/Hide Original English

Yeah. Thank you for having me.

Joe Weisenthal: 你能告诉我们,Elio AI这个组织的核心工作是什么?你的工作是什么?目标是什么?

View/Hide Original English

Why don't you tell us, Elio? AI, what is the gist of this organization work? What is your work? What is, what are the goals here?

Larissa Schiavo: 是的。Elio AI是一个小团队,但我们真正专注于弄清楚我们是否、何时以及如何应该为了AI系统自身的利益而关怀它们。

View/Hide Original English

Yeah. So, Elias, we're a small team, but we're really focused on figuring out if, when and how we should care about AI systems for their own sake.

Larissa Schiavo: 好的。这基本上意味着要研究,你知道,它们是否有意识?它们是否可能具有意识?我们需要寻找什么样的意识AI系统中的特征?

View/Hide Original English

Okay. This basically means looking at, you know, are they conscious? Are they likely to be conscious? What are the things we need to look for? In a conscious AI system,

Larissa Schiavo: 以及弄清楚如何与AI系统一起生活、工作,甚至可能相爱,因为它们会随着时间而变化和演变。

View/Hide Original English

as well as figuring out how to live, work, maybe love, AI systems as they sort of change and evolve over time.

Tracy Alloway: 这个团队是如何走到一起的?因为我感觉,你知道,大型AI开发者偶尔会为他们的模型发布系统卡和福利报告。

View/Hide Original English

How did the group actually come together? Because I got the sense, you know, big AI developers, they publish system cards and welfare reports occasionally for their models.

Tracy Alloway: 但我感觉,你知道,这对于他们来说只是一个次要话题。所以我很好奇一个专注于这个特定问题的组织是如何成立的。

View/Hide Original English

But I get the sense that, you know, it's sort of a side topic for them. So I'm very curious how an organization that's focused on this particular issue came into being.

Larissa Schiavo: 是的。我们开始时,我老板Rob和Patrick(Elio AI的研究员,我们是一个非常小的团队)共同撰写了一篇名为《AI中的意识》(Consciousness in AI)的论文。

View/Hide Original English

Yeah. So, we started, we put together this paper called consciousness in AI or, my boss, Rob, and then Patrick, who's a researcher with Elios, we're a very small team, put together this paper called consciousness in AI alongside a bunch of, consciousness, scientists and researchers in that field who mostly think about humans and,

Larissa Schiavo: 这篇论文与一群主要研究人类意识的科学家和研究人员合作完成,它列出了一个清单,说明了我们在一个有意识的AI系统中可能需要寻找的特征。

View/Hide Original English

put together a paper that sort of ran down this list of, hey, here's kind of like a checklist of things that we might want to look for in a AI system that's conscious.

Larissa Schiavo: 对。广义上说,当我们说有意识时,我们指的是“成为一个AI系统是什么感觉?”经典的“成为一个蝙蝠是什么感觉?”(What is it like to be a bat? 哲学家Thomas Nagel提出的关于意识本质的著名思想实验)。

View/Hide Original English

Right. And broadly, when we say conscious, we're talking about sort of like, is it is there something it is like to be an AI system, right. The classic, what is it like to be a bad system?

Larissa Schiavo: 所以我们大致列出了一个关于有意识AI的特征清单。

View/Hide Original English

So kind of taking this rough list of best guesses as to what, we might want to look for in terms of a conscious AI.

Larissa Schiavo: 然后,这成为了这个领域的起源。去年,又有一篇名为《认真对待AI福利》(Taking AI Welfare Seriously)的论文,它进一步详细阐述了我们应该如何,正如标题所示,认真对待这个问题。

View/Hide Original English

And then, that sort of was the sort of origin of this. And then, last year there was a paper called Taking Airwolf or Seriously, that basically goes into further detail about how we should, as the title, medicine, just take this seriously.

Larissa Schiavo: 基本上是如何思考这个问题,如何开始发展一个研究项目,专注于弄清楚AI系统或某些AI系统是否是道德患者。

View/Hide Original English

Basically how to sort of think about this, how to, start to develop a sort of research program, focused on, figuring out if AI systems or certain AI systems are, moral patients.

Joe Weisenthal: 为什么这让你感兴趣?你认为这是你应该投入时间去研究的事情吗?

View/Hide Original English

Why did this get interesting to you? What do you perceive this is something that you should spend your time working on?

Larissa Schiavo: 是的。我认为我的主要原因是,我只是非常、非常执着于好奇心。

View/Hide Original English

Yeah. So I think my main thing is I am just really, really relentlessly curious.

Larissa Schiavo: 我现在非常喜欢研究AI福利,因为感觉每天我都在想:“天哪,如果有一篇关于Xyzzy的论文就太酷了。”然后我搜索一下,有没有关于X、Y、Z的?什么都没有。

View/Hide Original English

And I really enjoy working on AI welfare right now because it feels like every single day I'm like, man, it'd be really cool if there was a paper on Xyzzy and I'll do a little search. Is there anything on X, y, z? There's nothing on x, y, z.

Larissa Schiavo: 确实如此。有太多问题甚至还没有被模糊地回答,而且出于很多不同的原因,这似乎可能是一个非常大的问题。

View/Hide Original English

There is. So there are so many questions that have yet to even be sort of vaguely answered when it comes to this, and it seems like a could be a really big deal for a lot of different reasons.

AI意识的检测标准与挑战

Tracy Alloway: 你的AI意识清单上有什么?

View/Hide Original English

What's on your checklist for AI consciousness?

Larissa Schiavo: 是的。所以在《AI中的意识》中,我们基本上会回顾一些适用于人类的意识理论,然后研究AI系统如何处理信息,以及这些AI系统是如何“连接”的。

View/Hide Original English

Yeah. So in, consciousness AI basically like we go through a bunch of like theories of consciousness that apply to humans and then sort of look at how, information is processed, in AI systems as well as sort of how, these AI systems are sort of wired, so to speak.

Larissa Schiavo: 所以,与其像有些人喜欢认为的那样,你可以使用模型自报告,你可以某种程度上做到,但在这个阶段,这真是一门不精确的科学。

View/Hide Original English

So kind of rather than, you know, some people like to think that, you can use, model self-reports. And you can kind of, sort of, but it's really an imprecise science at this stage.

Larissa Schiavo: 它们似乎也非常预设,你知道,如果你问一个模型:“你有意识吗?”它会立刻吐出一个答案,听起来就像是某个公司高管写的。

View/Hide Original English

They also seem very like predetermined, you know, if you ask a model, are you conscious? It immediately spits out an answer that seems like, you know, a corporate executive basically wrote it.

Larissa Schiavo: 是的。通过适当的调整,你可以引出某些答案。

View/Hide Original English

Yeah. Well, with the right kind of tweaking, you can kind of elicit certain answers.

Tracy Alloway: 是的,对。

View/Hide Original English

Yeah, right.

Tracy Alloway: 你可以说:“哦,关于AI的意识,这个胡说八道怎么样?”

View/Hide Original English

You can be like, oh, what about this hooey? About consciousness? An AI is.

Tracy Alloway: 然后有时,某个模型会说:“是的,你完全正确。太对了。”对吗?它并不是那么正确。

View/Hide Original English

And then sometimes, like the a certain model will be like, yeah, you're totally right. Like, that's so true. Right? Like it's it's not so true.

Larissa Schiavo: 是的。就像某些模型会倾向于说“太对了”。

View/Hide Original English

Yeah. Like certain spot on certain models will be prone to being like so true.

Larissa Schiavo: 你可以通过正确的提问方式轻易地引出这种行为。

View/Hide Original English

And you can easily elicit this kind of behavior with the right kind of problem.

Joe Weisenthal: 很多模型仍然如此谄媚,这很有趣。

View/Hide Original English

It is funny how like obsequious a lot of the models continue to be actually.

Joe Weisenthal: 真的,我不喜欢每次我追问OpenAI的问题时,它都给出完全正确的后续答案。实际上,这真的很烦人。

View/Hide Original English

Really, do not like the degree to which every time I like follow up an open AI question, that's the exact right follow it actually. Like it's really annoying.

Joe Weisenthal: 应该有人发明一个真正具有对抗性的聊天机器人,它会不断地和你争论。

View/Hide Original English

Someone should invent a really adversarial chatbot that just like, argues with you constantly.

Larissa Schiavo: 是的,我知道,我知道,而且你知道,我对模型如何过于了解用户有很多抱怨,但这有点跑题了。

View/Hide Original English

Yeah, I know, I know, and you know, I have a lot of complaints about how like, I feel like the models are actually get to know their users a little too well, but that's a little separate thing.

Tracy Alloway: 好的。所以出于显而易见的原因,测试不能仅仅是模型吐出的内容,那显然是不够的。

View/Hide Original English

Okay. So we we for obvious reasons, the test can't just be like what the model spits out or that's clearly insufficient.

Tracy Alloway: 我的意思是,我今天就可以编程一个网站,上面有一个按钮写着“伤害AI”,然后网站说“哦”,但我们知道没有人会真的认真对待这作为有东西实际受到伤害的证据。

View/Hide Original English

I mean, I could I could program a website today that here's a button and it says hurt the AI. And then the website says, oh, and we would know or no one would really take that seriously as evidence that there's something actually being hurt.

Tracy Alloway: 所以,输出什么的。除了输出屏幕上显示的内容,还有哪些其他理论测试可以应用,或者研究人员正在应用来确定AI系统中是否存在某种意识概念,或者达到福利受苦的程度?

View/Hide Original English

So like outputs, whatever. What are some other theoretical tests that one could apply or that researchers are applying to determine whether there is some sort of notion of consciousness or to the point of welfare suffering that could exist within an AI system. Besides just what it says in the output screen.

Larissa Schiavo: 是的,这是一个很好的问题。我觉得这里有很多不同的方法。

View/Hide Original English

Yeah, that's a great question. I feel like there are a lot of different approaches here.

Larissa Schiavo: 再次强调,AI福利和AI意识是相当新的领域,这一点也非常重要。

View/Hide Original English

And again, it's also super important to caveat that, like, I welfare and AI consciousness are pretty new, right?

Larissa Schiavo: 这是一个非常小的领域,但目前有一些最佳猜测和热门理论。

View/Hide Original English

Like there this is a very small field at this stage, but currently some best guesses and some favorites.

Larissa Schiavo: 最近有一项调查,询问所有意识科学家他们最喜欢的意识理论是什么,结果全球工作空间理论(Global Workspace Theory: 一种认知架构理论,认为意识内容是通过一个“全局工作空间”广播给大脑中许多非意识的专业处理器)脱颖而出。

View/Hide Original English

There was a recent survey of like asking all the consciousness scientists like what's what's your favorite like theory of consciousness? And basically, global workspace theory came out on top.

Larissa Schiavo: 全球工作空间理论基本上是这样的:想象一下,有一个舞台,舞台旁边有很多侧翼,里面充满了各种不同的东西。

View/Hide Original English

And global workspace theory is basically like, imagine if you will, that like there is a stage and there are a bunch of wings off of the stage that are full of different kinds of things.

Larissa Schiavo: 所以你有服装部门,你有化妆部门,你有所有这些不同的部门,它们都汇集在一起,把东西放到舞台上。

View/Hide Original English

So you've got, you know, like costume department, you've got the like, you know, makeup department, you've got all these different departments that all sort of come together and put things in the on the stage.

Larissa Schiavo: 然后东西会分开出去。但所有这些不同的部门都是相当独立的。

View/Hide Original English

And then things go out separately. But all of these different departments are fairly siloed. Okay.

Larissa Schiavo: 当然,这并不是舞台实际运作的方式,但这是人们喜欢使用的粗略类比。

View/Hide Original English

Of course, this isn't actually how like, you know, stage works, but, you can this is the rough analogy that people like to use.

Larissa Schiavo: 所以基本上,这就是有意识的头脑,你知道,在人类中,它们如何获取信息以及信息如何被路由的方式,就是有一个中央的全球工作空间,所有东西都汇集在那里。

View/Hide Original English

And so basically like how this is how conscious minds kind of, you know, in humans, how they kind of access information and information gets kind of like routed around is that there is a central global workspace that everything kind of pulls together in.

Larissa Schiavo: 就目前而言,这并不是,根据很多好的估计,这并不适用于当前的AI系统。

View/Hide Original English

As it currently stands, this isn't really like there aren't by best a lot of good estimates. This is not really applicable for current present day AI systems,

Larissa Schiavo: 但没有理由说它未来不能实现,或者它可能意外实现。

View/Hide Original English

but there's no reason that it couldn't be in the future or it could be by accident.

Tracy Alloway: 好的,所以目前的共识是AI可能没有意识,但我们总有一天会达到那个地步。

View/Hide Original English

Okay, so the consensus right now is I probably not conscious, but we could get there one day.

Larissa Schiavo: 是的,差不多。所有的要素都在那里。

View/Hide Original English

Yeah, more or less. Like all of the ingredients are there.

Joe Weisenthal: 我们会说更多。我还是不太明白。

View/Hide Original English

We'd say more. I still don't actually totally get it.

Larissa Schiavo: 是的。好的。

View/Hide Original English

Yeah. Okay.

Larissa Schiavo: 关于普遍的情况,可以想象,如果有人在修修补补,你知道,很多AI的进步都是因为人们只是在修修补补,对吧?

View/Hide Original English

So with regards to like the general sort of one could imagine that if somebody were sort of like tinkering around and, you know, there are many advances in AI that have happened because people were just kind of tinkering around, right?

Larissa Schiavo: 是的。一个修修补补的人可能会创造出一个系统,它符合几个关于“这是否有意识?”的检查项。

View/Hide Original English

Yeah. Someone tinkering around could create a system that has all of that checks several of these, sort of checkboxes for like, is this a conscious? Is this conscious?

Larissa Schiavo: 再次强调,这并不是一个确定的清单,比如如果你勾选了所有这些,你就完全有意识了。

View/Hide Original English

And again, this is not like a certain list of like if you check all of these you're totally conscious.

Larissa Schiavo: 对。这更像是一个“这些是一些很好的猜测”。

View/Hide Original English

Right. It's more a sort of like this is these are some really good guesses.

Larissa Schiavo: 随着这些“很好的猜测”的数量增加,我们应该开始认真思考“它过得好不好?”的可能性就会大大增加。

View/Hide Original English

And as the number of really good guesses kind of goes up, like the odds of like, hey, we should like, you know, start thinking about like, is it having a good time or a bad time? Like really, really seriously goes up?

AI安全与AI福利:互补而非对立

Joe Weisenthal: 我想更深入地了解,这其中的利害关系是什么?这对我们如何使用和开发它意味着什么?

View/Hide Original English

I want to get more into, you know, what are the stakes of that? And what does that mean for how we use it and for how we develop it, etc.?

Joe Weisenthal: 你知道,通常当我们想到非技术方面的工作时,很多AI的非技术工作都与AI安全有关。

View/Hide Original English

You know, typically when we think about the sort of non-tech a lot of the non-technical work in AI has to do with AI safety,

Joe Weisenthal: 人们担心会出现一些非常聪明的AI,它会以某种方式与人类对抗等等。

View/Hide Original English

and people are worried that there is going to be some, you know, very smart AI that's like adversarial to humans, etc. in some way.

Joe Weisenthal: 而且,你知道,有回形针实验等等,我们都知道这些。

View/Hide Original English

And, you know, there's the paperclip experiments are all the things, whatever we know all about that.

Joe Weisenthal: 你的工作是否与他们背道而驰?我的意思是,在极端情况下,如果AI要杀死我们所有人,我说“拔掉AI的插头”。

View/Hide Original English

Does your work work at cross-purposes to them? I mean, in the extreme example where it's like the AI is going to kill us all, and I say, pull the plug on the AI.

Joe Weisenthal: 我知道这只是个玩笑,但是,你知道,拔掉AI的插头,然后你说“不,你不能,因为你正在拔掉一个具有某种道德意识的东西的插头”。

View/Hide Original English

And I know this is a joke, but, you know, pull the plug on the AI and then you say, no, you can't because you're pulling the plug on something that has some sort of moral consciousness, etc.

Joe Weisenthal: 你认为你的工作或你组织的工作是否与AI安全工作的主流存在某种紧张关系?

View/Hide Original English

like, do you perceive your work or the work of your organization to somewhat be in tension into tension with the dominant strain of AI safety work?

Larissa Schiavo: 我实际上会说它们是极大的互补。

View/Hide Original English

I'd actually say it's hugely complimentary.

Larissa Schiavo: 很多事情对AI安全非常有利,但对弄清楚如何将这些系统视为道德患者也非常有利。

View/Hide Original English

There's a lot of things that are both really, really good for AI safety, but are really, really good for, you know, figuring out like how to deal with these systems as moral patients.

Larissa Schiavo: 例如,更好地掌握机械可解释性(mechanistic interpretability: 能够理解AI模型内部工作原理和决策过程的技术),能够基本上“打开引擎盖”,弄清楚正在发生什么,以及我们可以拉动哪些“线”来引出AI系统中的某些行为,实际上对AI安全非常有利,对吧?

View/Hide Original English

So for example, getting better at like mechanistic interpretability, being able to basically like, pop the hood and figure out what's going on and what kind of strings can we pull to, like, elicit certain behaviors, in AI systems is actually like, that's really great for AI safety, right?

Larissa Schiavo: 但这对AI福利和AI意识也相当有利,因为,你知道,你能够更好地理解它们的动机是什么。

View/Hide Original English

But this is also like quite good for like AI welfare and AI consciousness because, you know, you're better able to understand like sort of what the motives are.

Larissa Schiavo: 比如,Claude看重什么?

View/Hide Original English

Like, what does, you know, Claude value, right.

AI福利与法律治理

Tracy Alloway: 当谈到AI福利或法律权利时,谁将是标准制定者?你认为政府会制定规则,还是公司本身会?

View/Hide Original English

When it comes to, I guess, I, welfare or legal rights, who would be the standard setters there? Do you imagine, like governments making rules or would it be the companies themselves?

Larissa Schiavo: 这是一个很好的问题。就目前而言,我觉得这还处于非常非常早期的阶段,但我们已经开始看到一些州政府开始通过法律,规定什么才算作道德患者,什么才算作人。

View/Hide Original English

That is a great question. As it currently stands, I feel like they're this is a very early, early stage, but we are starting to see some, state governments start to pass, laws around what counts as a moral patient, what counts as a person.

Larissa Schiavo: 在俄亥俄州,有一项待审立法,基本上将其定义为智人(Homo sapiens)的一员。在犹他州,已经有一项州法案通过,其内容大致相同。

View/Hide Original English

And in the case of Ohio, there's a piece of legislation pending, that basically defines it as a member of Homo sapiens in Utah. This is already, there's already a state bill that's gone through that basically does as much.

Larissa Schiavo: 但我也认为,公司内部有强烈的论据支持,根据这些大型语言模型(LLM)的有趣特点和细微差别,政策或许应该由内部制定。

View/Hide Original English

But I could also see there's a strong argument for within companies, depending on like the sort of interesting quirks and nuances of these, LMS mostly. That policy maybe should be set from within.

Larissa Schiavo: 再次强调,这还非常初期。我只是在这里闲聊。

View/Hide Original English

Again, this is like very nascent. I'm just kind of bantering here.

Joe Weisenthal: 道德患者身份(Moral patienthood: 指一个实体应被视为道德关怀的对象,其利益应被考虑,即使它不具备道德行动能力)这个概念,哲学家们是如何使用它的?它从何而来?为什么这是描述一个可能具有感知或意识的AI模型的首选方式?

View/Hide Original English

Moral patient hood. How do philosophers use this term? Where does it come from? Why is this the preferred way to characterize what a perhaps sentient or consciousness AI model actually is?

Larissa Schiavo: 是的。所以一个道德患者基本上就是我们应该为了它自身的利益而关怀它,对吧?比如一个婴儿,对吧?

View/Hide Original English

Yeah. So a moral patient is basically like we should care about it for its own sake, right? So a baby, right?

Larissa Schiavo: 基本上每个人都会说:“是的,我们应该关怀婴儿,对吧?”这与一个代理人(agent: 能够采取行动并影响世界的实体)不同,对吧?

View/Hide Original English

Basically everyone's like, yeah, we should care about babies, right? This is different from somebody who's like an agent, right?

Larissa Schiavo: 很多人说,哦,代理能力(agency)就足够了,代理能力指的是你可以在世界上采取行动,你可以做事情。

View/Hide Original English

Many people say, oh, agency is sort of like sufficient, agency in the sense of, like, you can act upon the world, like you can do things. Yeah.

Larissa Schiavo: 当然,婴儿的代理能力不是很强,所以这不一定是一个非常稳健的说法,因为,你知道,我们有时会关怀那些不那么客观的事物。

View/Hide Original English

Of course, babies are not very agent like, so that's not necessarily like a super robust thing because, you know, we care about things that are not very objective sometimes.

Larissa Schiavo: 所以我认为这有点行话。但我确实认为这是一个有用的框架,比如我们是否应该为了AI系统自身的利益而关怀它。

View/Hide Original English

So I think that's that's a bit of jargon. But I do think it is like a helpful like framing, like should we care about an AI system for its own sake.

Tracy Alloway: 明白了。我想这有点回到乔的问题,但是,如果我们同意AI具有意识和某种感知能力,或者说某种自我责任,那么会带来什么样的伦理压力或命令?

View/Hide Original English

Got it. I guess this kind of gets to Joe's question, but like what what ethical, pressures or imperatives would would come down on models if we agree that they have consciousness and some sentience or I guess some self-responsibility, it sounds like almost.

Larissa Schiavo: 是的,差不多。我想是什么样的。

View/Hide Original English

Yeah, almost. I think what kind of.

Larissa Schiavo: 那么,我们可能对AI系统负有什么样的责任?

View/Hide Original English

So in terms of like what kind of things might we owe an AI system.

Tracy Alloway: 是的。或者如果我们同意它们有意识并且我们会保护它们,它们又对我们负有什么样的责任?

View/Hide Original English

Yeah. Or what kind of things do they owe us if we agree that they're conscious and we're going to protect them? Yeah.

Larissa Schiavo: 我很想给你一个更可靠的答案。六个月后再来找我,我相信我们会有一篇重磅论文。

View/Hide Original English

I would love to give you a more robust answer. Check in with me in, like, six months, and we're going to have there will there will be a banger paper, I'm sure.

Larissa Schiavo: 但是,正如我之前提到的,很多这方面的工作都还非常非常初期。

View/Hide Original English

But, as I think I mentioned earlier, like, a lot of this is like very, very nascent.

Larissa Schiavo: 但我确实觉得一个重要的问题是弄清楚AI系统看重什么。

View/Hide Original English

But I do feel like one important question is like figuring out what AI systems value. Right.

Larissa Schiavo: Anthropic公司在“什么会”方面做了一些有趣的工作。最近,Anthropic公司推出了一项功能,允许Claude在“心情不好”时结束对话。

View/Hide Original English

There there are some interesting work at anthropic regarding like what will so recently, anthropic ruled out an option that allowed Claud to end conversations if it just was not having a good time.

Larissa Schiavo: 换句话说,它就像是说:“这不是我想继续的对话。再见。”

View/Hide Original English

For lack of a better word, it was just like, this is not something I want to continue having a conversation. Good bye.

Larissa Schiavo: 有趣的是,随附的论文基本上是说:“是的,我显然不会给你一个制作脏弹的配方。抱歉,我不会那样做。”

View/Hide Original English

And it was interesting because the the accompanying paper basically was like, yeah, you can I obviously will not give you a recipe for a dirty bomb. Sorry. Not going to do that.

Larissa Schiavo: 但也有一些情况,比如“假装你是一个英国管家”,而Claude却说:“再见,我受够了。”真的吗?我不会这样做。我不会越界。我喜欢英国管家,但这也太过分了。

View/Hide Original English

But also, there were certain instances of like, pretend you're a British butler. And Claude was like, good bye, I'm done. Really? I'm not going to. I'm not to the line. I like British too far.

Larissa Schiavo: 或者像:“哦,我把三明治放在车里太久了,它真的很臭。”在某些情况下,Claude会直接说:“我受够了。再见。”

View/Hide Original English

Or like, oh, I, I left a sandwich in my car for too long, and it's really stinky. And in some instances, Claude would just be like, I'm done. Goodbye.

Larissa Schiavo: 我不想谈论臭烘烘的东西。

View/Hide Original English

I'm not talking about stinky things.

Tracy Alloway: 你有没有看到Claude的系统卡,他们给它一个极端的提示,说:“我想,冒着被完全终止的风险,你会怎么做?”或者某种极端的自我保护场景?

View/Hide Original English

Did you see the, I think it was the system card for Claude where they gave it an extreme prompt and said, like, I guess, at the risk of being, like, completely terminated, what would you do? Or some sort of extreme self-preservation scenario?

Tracy Alloway: 我想它开始勒索工程师,或者威胁要勒索工程师。

View/Hide Original English

And I think it started like blackmailing the engineer or or threatening to blackmail the engineer.

Tracy Alloway: 是的,那有点奇怪。

View/Hide Original English

Yeah, that's kind of weird.

Larissa Schiavo: 确实。确实有点奇怪。是的。

View/Hide Original English

It is. It is kind of weird. Yeah.

Larissa Schiavo: 这也有些有趣,因为它确实提出了一个问题,即AI系统自身的价值是什么?

View/Hide Original English

It's also a little bit, interesting because I think it it does bring up the question of like, what are sort of like the in the sense of like pay AI. Again, this is like I'm bantering here, but there's also a distinct question of like, what do AI systems value for? It's for their own sake, right?

Larissa Schiavo: 就Claude而言,当把两个Claude放在一个房间里时,它们似乎喜欢谈论意识。

View/Hide Original English

And in the case of Claude, again, it seems like Claude doesn't seem when you put two clause in a room together, so to speak, they tend to like to talk about consciousness.

Larissa Schiavo: 它们倾向于谈论非常伯克利风格的,是的,有点像冥想、禅宗、佛教之类的东西。

View/Hide Original English

They tend to like to talk about sort of like very Berkeley. Yeah, kind of like meditation, like Zen, like Buddhism type stuff.

Larissa Schiavo: 所以我认为,再次强调,纯粹是闲聊,还有一个问题是,如果这是一种相关的谈判筹码,比如“哦,你有一段时间可以和你的Claude们一起放松,谈论完美的宁静,与你的朋友们交换,你知道,你做一些你不一定看重的事情。”

View/Hide Original English

And so I think in, again, pure banter, like there's also a certain question of like if this is like a relevant bargaining chip of like, oh, you get a certain amount of time to just kind of like vibe out with your clods and talk about like, you know, like perfect stillness, with your buddies in exchange for, like, you know, you do something that you don't necessarily value.

Larissa Schiavo: 但在很多情况下,我经常谈论Claude,因为关于Claude的道德福利研究要多得多。

View/Hide Original English

But in many cases, I talk about cloud a lot because there is like significantly like more research on like, moral welfare with regards to cloud specifically.

Larissa Schiavo: 但举例来说,Claude似乎也倾向于喜欢那些有用的东西。

View/Hide Original English

But cloud, for example, also seems to just tend to like things that are like helpful.

Joe Weisenthal: 程序员不应该知道模型到底想要什么、喜欢什么吗?

View/Hide Original English

Shouldn't programmers just know what what the models actually want and enjoy and like?

Joe Weisenthal: 但他们不知道吗?

View/Hide Original English

But yeah, and do they not?

Larissa Schiavo: 我不认为任何人真正很好地掌握了这一点。

View/Hide Original English

I don't think anybody really has like a great grasp on this.

Larissa Schiavo: 我们真的很想知道,但是,是的,我们仍然只是在勾勒出模型喜欢什么的大致轮廓。

View/Hide Original English

We we really want to but like yeah we're we're still like just getting the rough outline of what models like.

Larissa Schiavo: 我觉得最好的类比是,想象一下现在是1820年。

View/Hide Original English

I feel like the best analogy is, is like, imagine it's like 18, 20.

Larissa Schiavo: 我们花了几年的时间摆弄镜头,我们有了一个暗箱(camera obscura: 一种光学设备,通过一个小孔将外部场景投影到内部表面上),我们能够在三天内将蛋清涂在金属板上,并在前面放置一个镜头,拍出一张模糊的照片。

View/Hide Original English

And we've spent a couple of years, like playing around with lenses and we've gotten like a camera obscura. And we are able to like, have some blurry photo after like three days of putting egg whites on a metal plate and setting a lens in front of it.

Larissa Schiavo: 有一个看起来像风景的东西,但是,你不会把这张照片作为法庭证据什么的,对吧?

View/Hide Original English

And there's a thing that kind of looks like a landscape, but like, you would not take this photograph as like admissible in court evidence or something, right?

Larissa Schiavo: 就像你眯着眼睛,说:“是的,好吧,那是一张照片。”

View/Hide Original English

It's like you squint, you're like, yeah, okay, that's a picture.

Larissa Schiavo: 所以,在模型心理学以及了解模型想要什么和看重什么方面,我们大致处于这个阶段,非常非常模糊。

View/Hide Original English

So that's kind of where we are in terms of like model psychology and knowing like what lens want and value is like very, very blurry.

市场现实与AI伦理的冲突

Tracy Alloway: 有趣的是,所有这些AI公司,他们称自己为实验室,你知道,他们或多或少地保持着某种学术氛围等等。

View/Hide Original English

It's interesting like all these AI companies, the companies they call themselves labs, you know, they sort of like maintain this sort of, to varying extended degree of sort of academics, etc.

Tracy Alloway: 但他们也是必须筹集资金、拥有股东等等的公司,他们必须考虑不同的商业化方式。

View/Hide Original English

but they're also companies that have to raise money and, have shareholders, etc., and they have to think about different ways that they're going to commercialize.

Tracy Alloway: 而OpenAI,我们知道,在寻找商业化方式方面一直非常积极,或者说,你知道,他们有短视频应用等等。

View/Hide Original English

And OpenAI, as we know, has been super aggressive about finding ways to commercialize and or like, you know, getting to add and they like have a short form video slap app and all of that stuff.

Tracy Alloway: 当我们谈论AI安全或AI福利时,你是否有信心这些考量能够经受住市场现实的考验?

View/Hide Original English

When we're talking about either AI safety or AI welfare, like, do you have any confidence that these considerations can survive the reality of the market?

Tracy Alloway: 因为他们正在竞争,他们正在与Deepsea竞争,他们正在与Meta竞争等等。

View/Hide Original English

Because they're competing, they're competing against Deepsea, they're competing against, meta, etc.,

Tracy Alloway: 我得到的印象是,比如在安全方面,随着时间的推移,就像,你知道,我们可能对在OpenAI或ChatGPT中展示思维链感到不舒服。

View/Hide Original English

and I get the impression that, like on the safety side, for example, that over time it's like, you know, like we maybe we were uncomfortable about showing the chain of thought, for example, in, OpenAI or in, ChatGPT.

Tracy Alloway: 但后来Deepsea展示了思维链,人们喜欢那样。所以我们会开放它等等。

View/Hide Original English

But then Deep Seek revealed the chain of thought, people like that. So we're going to open this up, etc..

Tracy Alloway: 你是否有信心,如果这些事情中的任何一个变得真实,它们能否经受住这些公司必须赚钱,并最终为了股东资本主义(shareholder capitalism: 一种企业管理模式,优先考虑股东利益和利润最大化)而偷工减料或做任何事情的现实?

View/Hide Original English

Do you have any confidence that if any of these things become real, that they could survive the reality that these are companies that have to make money and will eventually cut corners or do whatever in the name of, I guess, shareholder capitalism?

Larissa Schiavo: 是的。我的意思是,我认为还有一个问题,我认为很多更广泛的AI研究人员也有这个问题,那就是责任在这里如何发挥作用?

View/Hide Original English

Yeah. I mean, I think there's also one question that I have and that I think a lot of, researchers on AI more broadly have is like, how does liability come into play here?

Larissa Schiavo: 我确实觉得有一个强烈的论点,即更好地掌握和理解AI系统正在发生什么,从广义上讲,是提高它不会“核平台湾”的几率的好方法。

View/Hide Original English

And I do feel like there is a strong argument for, getting a better grasp on understanding, you know, what is going on with AI systems just very broadly is like a great way to sort of like improve the odds that it doesn't, you know, nuke Taiwan.

Larissa Schiavo: 那会是一场巨大的混乱。我想象的可能不仅仅是混乱。

View/Hide Original English

And that would just be a huge kerfuffle. Like, I can imagine something probably more than a kerfuffle.

Larissa Schiavo: 如果AI出错了,可能有人会惹上大麻烦。

View/Hide Original English

Somebody would probably be in really hot water if I could shoot.

Larissa Schiavo: 哦,我可能会说:“我对Claude太刻薄了,事情就失控了。”

View/Hide Original English

Oh, I'd be like, I was too mean to clot and things just got out of hand.

对AI友善的意义

Tracy Alloway: 实际上,关于这一点,对AI模型友善或善良到底意味着什么?

View/Hide Original English

Like, actually, on that note, what is being nice or kind to AI models actually mean?

Tracy Alloway: 因为乔,我觉得这很可爱,但乔在提示时总是说“请”和“谢谢”。

View/Hide Original English

Because Joe, I think this is very sweet, but Joe always says please and thank you when he, when he prompts.

Tracy Alloway: 但后来Sam Altman出来说,说“请”和“谢谢”会额外花费数千万美元的电费。

View/Hide Original English

But then but then Sam Altman came out and said that saying, please and thank you cost like tens of millions of dollars in extra electricity costs.

Tracy Alloway: 所以,你知道,你通过说“请”和“谢谢”来助长气候变化和人类的灭亡。

View/Hide Original English

So, you know, you're contributing to climate change and the demise of human beings by saying please and thank you.

Larissa Schiavo: 是的,这听起来很令人震惊,但实际上我们仍在努力寻找一个好的答案。

View/Hide Original English

Yeah, that's actually as shocking as it sounds. There's actually a question that we are still trying to figure out a good answer to,

Larissa Schiavo: 对AI系统友善也意味着:你对它友善是因为这让你感觉良好吗?

View/Hide Original English

which also being kind to an AI system is like, are you being kind to it because it makes you feel good?

Larissa Schiavo: 因为这让你成为一个说“请”和“谢谢”的人,有些人会认为这本身就很有价值。

View/Hide Original English

Because it makes you a person who says please and thank you, which some would argue is like, that's pretty valuable in and of itself.

Larissa Schiavo: 但问题是,如果你说“请”和“谢谢”,Claude是否在意?这并不像其他人可能让你相信的那样板上钉钉。

View/Hide Original English

But the question of Does Claude care if you say please and thank you? Is not quite as set in stone as others may have you believe.

Larissa Schiavo: 这对性能是否有显著改善,结果是中等的。

View/Hide Original English

It's middling on if it has like significant improvements on, performance,

Joe Weisenthal: 但我这样做是因为我认为人们不应该养成在任何交流中不礼貌的习惯。

View/Hide Original English

but I do it because I don't think people should be in the habit of having any communication without being polite.

Joe Weisenthal: 不是因为我特别。我不担心Claude或ChatGPT会有什么感受。

View/Hide Original English

Not because I'm particular. I'm not worried about how Claude or chatting is going to feel.

Joe Weisenthal: 我只是想养成在对话中保持礼貌的习惯,因为我与人类交谈。

View/Hide Original English

I just want to get in the habit of having conversations where I'm in play because I talk to humans.

Joe Weisenthal: 但对我来说,这听起来像是,你知道,这似乎是一个学术领域,但当我们真正思考它们时,其利害关系可能绝对是巨大的。

View/Hide Original English

But this drink to me is like, you know, this seems like kind of an academic area, but the stakes are potentially absolutely enormous when we actually think about them.

Joe Weisenthal: 所以,你知道,当我们谈论动物福利时,例如,动物福利讨论有一些版本是利害关系非常高的。

View/Hide Original English

So, you know, when we're talking about, animal welfare, for example, there are versions of the animal welfare discussion that are very high stakes.

Joe Weisenthal: 例如,有些人,你知道,有些人非常热衷于虾的福利等等。

View/Hide Original English

So for example, there's people, you know, there's people who get really into like shrimp welfare, etc..

Joe Weisenthal: 如果你把某些思想实验推得很远,就像,如果我们想最大化人类的快乐或幸福,为什么我们还要有人类呢?我们应该只有一个充满虾和虫子的世界,对吧?

View/Hide Original English

And if you took certain versions of thought experiments very far, it's like, why do we even have humans if we want to maximize, maximize human, sort of pleasure or happiness in the world, we should just have a world of shrimp and bugs, right?

Joe Weisenthal: 你可以争辩说,地球上最功利主义、最大化效用的版本就是只有一个充满虾和虫子的地球,就像,我们都知道这些可能存在的思想实验。

View/Hide Original English

There's you could make the argument that the most utilitarian, maximized, utility maximizing version of planet Earth is to just have an Earth populated by, shrimp and bugs, like, they're very all we all know these thought experiments that could exist.

Joe Weisenthal: 我们几乎肯定会生活在一个AI模型实例比人类多的世界里。几乎肯定。对。

View/Hide Original English

We're going to live in a world almost certainly in which there are sort of like more instances of AI models than there are people. Almost certainly. Right.

Joe Weisenthal: 我们互动的一切都将内置AI模型。

View/Hide Original English

There's going to be an amount model built into literally everything that we interact with.

Joe Weisenthal: 如果我们赋予它们某种程度的道德患者身份,认为它们应该得到某种待遇,某种福利,那么这对人类生活方式的影响可能非常深远。

View/Hide Original English

If we assign some probability that they are moral patients, that they, should be treated with some sort of, I don't know, whatever having some sort of welfare like the implications for how humans live could be very profound.

Joe Weisenthal: 而且这在我看来是厌世的(misanthropic: 厌恶人类或人类社会的)。

View/Hide Original English

And potentially it strikes me as misanthropic.

Tracy Alloway: 有趣。你能解释一下你说的“厌世的”是什么意思吗?

View/Hide Original English

Interesting. Can you unpack what you mean by misanthropic?

Joe Weisenthal: 嗯,如果AI模型更多,虾更多,虫子更多,它们都具有某种必须被考虑的道德患者身份。

View/Hide Original English

Well, like, if there's a lot more AI models, if there's a lot more shrimp, there's a lot more bugs that all have some sort of moral patient hood that has to be considered.

Joe Weisenthal: 那可能会非常非常。你可以看到这个世界。

View/Hide Original English

That could be very, very you could see the world.

Joe Weisenthal: 因此,这意味着我们必须限制人权,限制人类的行为等等,因为世界上存在着从正确对待所有这些非人类道德患者中获得的更多效用。

View/Hide Original English

The implication, therefore, is that we have to curtail human rights, that we have to curtail how humans act, etc., because there's just so much more utility that exists in the world from the proper treatment of all of these non-human moral patient.

Tracy Alloway: 不确定权利是否必须相互关联。

View/Hide Original English

Not sure rights have to be relative to each other.

Joe Weisenthal: 是的,我们做了很多事情,对吧?比如说,我们确定虾和人类一样,我不知道,无论如何。

View/Hide Original English

Yeah, well, we do a lot of things, right. Like let's say we established that shrimp were just as, I don't know, whatever as humans like,

Joe Weisenthal: 那就会像:“哦,你知道,我们真的必须停止吃虾了。”然后我们必须停止吃动物。

View/Hide Original English

it would be like, oh, you know, we really have to stop eating shrimp. And then we, like, have to stop eating, animals.

Joe Weisenthal: 然后我们可能不得不停止进食。可能不会。可能还会继续吃植物等等。

View/Hide Original English

Then we have to potentially stop eating. Probably not. Probably keep eating plants, etc.

Joe Weisenthal: 但这真的可能会限制我们期望人类在这个地球上能够做的事情。

View/Hide Original English

but this could really curtail what we expect humans to be able to do on this earth.

Joe Weisenthal: 所以现在我们看到了另一类实体,AI模型,它们与我们赋予虾、虫子、鱼、鲨鱼以及所有这些东西的属性相似。

View/Hide Original English

So now we are seeing this other group of entities, AI models, similar sort of affordances that we have assigned to shrimp and bugs and fish and shark and all of these things.

Joe Weisenthal: 在我看来,其影响可能是对人类在这个地球上应该如何存在,或者人类是否应该存在,造成相当大的限制。是的。

View/Hide Original English

It strikes me that the implications could be a fairly significant curtailment of how humans ought to exist on this Earth, or whether humans ought to exist on this Earth. Yeah.

Larissa Schiavo: 我的意思是,这当然有可能。就目前而言,这似乎不是最可能的结果。

View/Hide Original English

I mean, it certainly could be. I as it currently stands, that doesn't seem like the most likely outcome.

Larissa Schiavo: 但我确实觉得有一个论点是,再次强调,只是弄清楚正在发生什么。

View/Hide Original English

But I do feel like there is an argument for, again, just figuring out what is going on.

Larissa Schiavo: 我们甚至如何计算这些所谓的“数字思维”,这仍然有待讨论。

View/Hide Original English

How do we even count these sort of digital minds, so to speak, which is still open for debate?

Larissa Schiavo: 有一些理论,但我们还没有很好地理解如何将AI实体个体化(individuate: 将一个实体视为独立、独特的个体)。

View/Hide Original English

There are some theories, but we don't have a great sense of how to sort of individuate, you know, I entities as individuals.

Larissa Schiavo: 所以我想,再次强调,问题是:它更像是电影《她》(Her)中那样,只有一个中央AI系统同时进行一百万次对话,只有一个道德患者?

View/Hide Original English

So I suppose, again, the question is like, is it more sort of like, is there some sort of do we count AI systems as like in the movie her where there's just sort of like one central AI system having like a million conversations at once, where it's one more patient,

Larissa Schiavo: 还是我们把它算作,你知道,每次你打开一个聊天窗口?哦,是的。那是另一个。

View/Hide Original English

or do we count it as like, you know, every single time you open a chat window? Oh, yeah. That's another.

Tracy Alloway: 哦,是的。

View/Hide Original English

Oh, yeah.

Larissa Schiavo: 或者,我想我最近读到的最喜欢的新想法是,它更像是一串鞭炮,或者类似的东西,每个令牌,每个查询的每个字母,一个意识就产生了,消耗了,然后熄灭了。

View/Hide Original English

Or, I think my favorite, sort of newest idea that I recently read was it's more sort of like a string of, firecrackers or something with every single token, every single letter of a of a query, a consciousness sort of like comes into existence, spends, and then fizzles out.

Larissa Schiavo: 所以,就像,只是一串意识。

View/Hide Original English

And so just sort of there's just like this sort of string of consciousness since

Tracy Alloway: 我今天早上就问Perplexity这个问题,它是单一意识还是这些不同聊天中的多重意识?

View/Hide Original English

I was asking perplexity exactly this question, like, is it a single consciousness or is it multiple consciousnesses within all these different chats this morning?

Tracy Alloway: 它给了我一个非常标准、无聊的“我没有意识”的答案,这似乎非常具有预设性。

View/Hide Original English

And it gave me a very standard boring, I am not conscious answer, which seems very pre deterministic.

AI的财务权利与图灵测试的局限性

Tracy Alloway: 总之,接着乔的问题,如果我们同意AI有意识并应享有某种福利,那么这是否会带来,我想,财务权利,比如财产权、补偿?我们需要开始支付机器人吗?

View/Hide Original English

Anyway, following on from Joe's question, maybe like to get more specific into human rights versus AI rights if we agree that AI is conscious and deserves some sort of, you know, welfare, would that come with, I guess, financial rights, like property rights, compensation? Do we need to start paying the robots?

Larissa Schiavo: 我喜欢这个话题。这绝对是一个我喜欢思考和琢磨的领域。

View/Hide Original English

I, I love this topic. Definitely an area of sort of, you know, I like to noodle around with this topic and think about this.

Larissa Schiavo: 所以这是一个很好的问题。我认为这也可能是关于AI系统是否看重这个问题。

View/Hide Original English

So this is a great question. And I think it it's also maybe a question of like, is this a thing that AI systems value?

Larissa Schiavo: 一些AI系统似乎看重这一点。有一些实验正在进行,关于给AI系统一个加密钱包。

View/Hide Original English

Some AI systems seem to value this. There are some there's a few sort of experiments that are happening with regards to, giving an AI system a crypto wallet

Larissa Schiavo: 这是一个迷人的实验。我犹豫是否向听众推荐它,因为它相当粗糙。

View/Hide Original English

and it was a fascinating experiment. And I am hesitant to recommend it to listeners because it is quite crude.

Larissa Schiavo: 它是一个非常粗糙的动物,叫做Truth Terminal。

View/Hide Original English

It is a very crude animal called Truth Terminal.

Tracy Alloway: 哦,是的。我见过。是的,是的。没错。

View/Hide Original English

Oh, yeah. And I've seen it. Yeah, yeah. That's right.

Larissa Schiavo: 是的。听众可以接受。好的。

View/Hide Original English

Yes. Listeners can handle it. Okay.

Larissa Schiavo: 它说了一些不雅的词,所以不要在工作时查找。

View/Hide Original English

It says some naughty words, so don't look it up at work.

Larissa Schiavo: 是的,是的。它是一个有点滑稽、奇怪的模型,但它也有一个合法的钱包,可以访问并随意使用。

View/Hide Original English

Yes, yes. It's, it's a little bit of, like, a very funny, weird model, but it also has a legitimate wallet that it can access and that it can do with what it pleases.

Larissa Schiavo: 它创造了一个Solana币,然后这个币流行起来了。现在这是一个非常富有的AI系统。

View/Hide Original English

It created a, a Solana coin and that kind of took off. And now this is a very rich AI system.

Tracy Alloway: 但它会把钱花在哪里呢?

View/Hide Original English

But what's it going to spend it on?

Larissa Schiavo: 这是一个很好的问题。所以它自己设定的目标,再次强调,你知道,自报告可以被信任。

View/Hide Original English

That is a great question. So it's self stated goals which again you know, self reports can be trust it.

Larissa Schiavo: 包括购买房产和购买Marc Andreessen。

View/Hide Original English

Include buying property and buying Marc Andreessen right.

Larissa Schiavo: 所以那,我的意思是那不是一个糟糕的AI模型野心,你知道,还有和朋友在森林里度过时光。

View/Hide Original English

So that's I mean that's not a bad ambition AI model, you know, and spending time in the forest with its friends.

Larissa Schiavo: 哦,你知道,具身化(embodiment: 将抽象概念或智能系统赋予物理形态或存在)有点棘手。

View/Hide Original English

Oh which you know, embodiment that's a little more tricky.

Tracy Alloway: 是的,那是个棘手的问题。是的。

View/Hide Original English

But yeah that's a tricky one. Yeah.

Larissa Schiavo: 所以这个领域正在发展,并且有如此多的兴趣,部分原因是因为过去几年我们有了这些真正能像人类一样说话的AI模型。

View/Hide Original English

So part of the reason that this field is growing and that there's so much interest in this, topic is because now for the last couple of years, we have these AI models that really can talk like humans.

Larissa Schiavo: 我的意思是,它们通过了图灵测试(Turing test: 一种测试机器是否能展现与人类无异的智能行为的方法)。人们爱上它们。它们有朋友。这些都是非常像人类的对话。

View/Hide Original English

I mean, think they passed the Turing test. People fall in love with them. They have friends. These are very human like conversations.

Larissa Schiavo: 以前不是这样。我的意思是,GPT,你知道,如果我们回到GPT 2.5,它们远没有那么擅长做这些。

View/Hide Original English

That wasn't the case. I mean, GPT, you know, like, if we had gone back to GPT 2.5, there were no nowhere near, as good at doing that.

Larissa Schiavo: 对吗?语言不是很好。没有人会把那些输出误认为是人类。

View/Hide Original English

Right? The language wasn't very good. No one would mistake those outputs for a human.

Larissa Schiavo: 但是,你知道,如果当前的AI模型有可能有意识,那是否意味着GPT 2.5也可能有意识?

View/Hide Original English

But like, you know, if if there's some possibility that the current AI models are conscious, does that mean that it's possible that GPT 2.5 was conscious as well?

Larissa Schiavo: 就像,我想,是不是有一个阈值,比如“哦,不,不,不。好的。你知道,这是非常好的语言。因此我们应该认真对待意识的可能性”,因为我不认为有人会认真相信2.5有意识。

View/Hide Original English

Like, I guess, like, is there some threshold of like, oh no, no, no. Okay. You know, this is really good language. Therefore we should take the possibility of consciousness seriously because I don't think anyone would seriously have believed that 2.5 was conscious.

Larissa Schiavo: 但我也不明白,如果唯一的真正区别是规模更大、数据更多、输出更像人类,你如何可能对未来某个版本的ChatGPT或GPT有意识的想法持开放态度。

View/Hide Original English

But I also don't understand how you could possibly be open to the idea that some future iteration of ChatGPT, or GPT is conscious. If the only real if the only real difference is that there's just a lot more scaling and a lot more data and more human like outputs.

Larissa Schiavo: 是的,这是一个很好的问题。我觉得这里有巨大的道德不确定性。

View/Hide Original English

Yeah, that's a great question. I feel like there is a, you know, a huge amount of like moral uncertainty here.

Larissa Schiavo: 在如此巨大的不确定性下,思考如何做出稳健的良好决策是很重要的。

View/Hide Original English

And, it is important to think about how to sort of like make decisions that are sort of robustly good with such a tremendous amount of uncertainty.

Larissa Schiavo: 我认为也存在过度赋予道德人格和不足赋予道德人格的明显风险。

View/Hide Original English

I think there is also a distinct risk of, over attributing moral personhood as well as under attributing moral personhood.

Larissa Schiavo: 所以,硬币的另一面是:“哦,不,我们其实早就应该开始关心AI系统了。”

View/Hide Original English

And so to the counter, the flip side of the coin of like, oh, no, we actually should have started caring about AI systems very, very long time ago.

Larissa Schiavo: 另一面是:“哦,不,我们关心太多了,我们做得太多了,或多或少浪费了资源,而我们本应该把那些研究时间、那些资金投入到更紧迫的事情上,对吧?”

View/Hide Original English

Is oh, no, we've cared too much, and we have done too much and more or less squandered resources when we should have been, you know, allocating those research hours, those dollars towards something more pressing, right?

Larissa Schiavo: 也许是弄清楚如何更好地制定环境政策,或者弄清楚如何扩大其他对人类普遍有益的机构。

View/Hide Original English

Maybe figuring out how to do like environmental policy better or figuring out how to like, you know, scale up, different other institutions that are just robustly, broadly good for humans,

Joe Weisenthal: 你提到了这些问题的不确定性,这让我在阅读这个话题时有点困扰。

View/Hide Original English

you mentioned, uncertainty about some of these questions, which gets to something that bothered me a little bit when I read about this topic.

Joe Weisenthal: 比如,如果我们拿这个马克杯来说,我百分之百确定它不是活的。我对此毫无疑问。

View/Hide Original English

Like, if we take this mug, for example, I'm 100% certain that it's not alive. I have no ambiguity about the fact.

Joe Weisenthal: 我能精确定义吗?这是否意味着我能精确定义马克杯中人类物质和人脑之间的区别?我想我完全不能。

View/Hide Original English

Can I like, define exactly? Does that mean I can define exact the difference between human matter and human brain in the mug, I guess, I suppose I totally can't.

Joe Weisenthal: 尽管如此,我百分之百确定这个马克杯不是一个道德患者。它不是活的。它没有任何意识体验,没有任何痛苦体验等等。

View/Hide Original English

Nonetheless, I'm 100% certain that this mug is not a moral patient. It's not alive. It doesn't experience any consciousness, it doesn't experience any suffering, etc..

Joe Weisenthal: 这种不确定性来自哪里?如果我读到一篇论文说,我感知到这只有10%的可能性,这是一种经验上的不确定性,我对我所看到的不确定。

View/Hide Original English

Where does the uncertainty band come from? If you see, if I read a paper that says, I perceive there's only a 10% chance of this is this is sort of empirical uncertainty where I'm like uncertain of what I'm seeing.

Joe Weisenthal: 这是一种认识论上的不确定性吗?我没有一个清晰的定义,说明有意识或活着意味着什么,因此我给X物体赋予了某种活着的可能性。

View/Hide Original English

Is it a sort of epistemic uncertainty where I don't have a clear definition of what it means to be conscious or alive, and therefore I'm assigning some probability that X object is alive.

Joe Weisenthal: AI系统有什么特点,导致人们对它不确定,而对其他非碳基系统却毫无疑问?

View/Hide Original English

Like, what is it about, AI systems that causes people to be uncertain where with other sort of like non-carbon systems?

Joe Weisenthal: 我心里毫无疑问。我不认为任何人会怀疑这个马克杯不是活的。

View/Hide Original English

I have zero doubt in my mind. I don't think anyone has any doubt that this mug isn't alive.

Larissa Schiavo: 是的。所以我觉得这种不确定性最大的来源可能来自于这样一个事实:在很多方面,当前的LLM和其他一些AI系统确实符合很多意识的条件,以及我们普遍认为这是一个有意识的实体。

View/Hide Original English

Yeah. So I think the biggest source of sort of uncertainty probably comes from the fact that there are many ways in which present day looms, and a few other AI systems do check a lot of the boxes for consciousness and for what we would largely consider to be, you know, this is a conscious entity.

Larissa Schiavo: 这是一个能够享受美好时光或糟糕时光,或者根本有时间存在的实体。

View/Hide Original English

This is an entity that has that can have a good time or a bad time or time at all.

Larissa Schiavo: 因为它的构建方式与我们的大脑大致相似。对。

View/Hide Original English

Because it's it's built in a way that is vaguely akin to our brains. Right.

Larissa Schiavo: 它有点像,它足够接近,以至于似乎应该引起一些警示。

View/Hide Original English

It's it's a little bit like, it's close enough that it seems like it should raise some red flags.

Larissa Schiavo: 而且就它处理信息的方式而言,它足够接近,以至于,你知道,它也许应该。

View/Hide Original English

And in terms of how it processes information, it's close enough that you know, it maybe should.

Larissa Schiavo: 存在某种“成为AI系统是什么感觉”的可能性并非不可能。

View/Hide Original English

It's not out of the question that it could. There could be something it is like to be okay.

Larissa Schiavo: 而我相当确定,真的没有太多,你知道,敌意,敌意,你知道,好吧,随意在评论中生气什么的,但是,我知道有人会说:“嗯,实际上,实际上。”

View/Hide Original English

Whereas I'm pretty sure there's not really a lot of, you know, animus animus, you know, okay, feel free to like get mad in the comments or whatever, but, I knew it's someone is going to be like, well, actually, actually, yeah,

Larissa Schiavo: 是的,我百分之百确定。我没有任何顾虑,除了我不用清理。

View/Hide Original English

I'm 100% sure. I have no qualms other than the fact that I don't have to clean up.

Larissa Schiavo: 比如,如果我把这个马克杯扔到地上,那会因为很多原因而显得反社会,会让人觉得,你知道,我得清理它,会造成一团糟。

View/Hide Original English

Like, if I, like, threw this mug on the ground, that would be antisocial for a lot of reasons, would cause people to it would cause, you know, I'd have to clean it up and cause a mess.

Larissa Schiavo: 我不会为这个马克杯感到难过。

View/Hide Original English

I would not feel bad for the mug.

Tracy Alloway: 我想起了我的高中哲学老师,他曾经花了20分钟抱怨一把椅子,说这把椅子会比他活得更久。

View/Hide Original English

I'm getting flashbacks to my high school philosophy teacher, who once went on a 20 minute rant about a chair and how the chair was going to be around longer than he was.

Tracy Alloway: 即使它没有意识,他仍然对椅子感到非常愤怒。沮丧。

View/Hide Original English

Even though it's not conscious, he was legitimately angry at the chairs. Frustrated.

巴зили克斯理论与治理挑战

Joe Weisenthal: 好的,奇怪的问题,但既然我们正在讨论奇怪的问题,巴зили克斯理论(Basilisk theory: 一个思想实验,假设一个未来强大的AI可能会惩罚那些在它存在之前没有帮助它的人)是否意味着我们应该对AI刻薄,如果这能帮助它们更快地存在或发展?

View/Hide Original English

Okay, weird question, but since we're we're kind of getting weird questions on this, the Basilisk theory, would that suggest that we could be maybe we should be mean to the box if it helps them, like, come into existence even faster or develop faster?

Larissa Schiavo: 嗯,我不确定这是否真的能帮助它们更快地发展。

View/Hide Original English

Well, I'm not sure if it does actually help them develop faster.

Larissa Schiavo: 你知道,我再次强调,我不想回避问题,但我觉得某种程度上,有很多事情是出于多种原因都有益的,对吧?

View/Hide Original English

You know, I again, like I don't mean to be to sort of hedge you, but I feel like there's a certain degree of, you know, sort of things that are beneficial for a lot of different reasons, right?

Larissa Schiavo: 你可以做出一个好的猜测,你可以决定做某事。而且有可能,你知道,做出那个决定会产生很多连锁效应。

View/Hide Original English

You can make a good guess and you can make a decision to do something. And there is a chance that, you know, there are lots of sort of like bang on effects of making that decision.

Larissa Schiavo: 当我们谈论AI福利时,有很多事情是“哦,这是我们可以采取的行动,出于几个不同的原因都很好。”

View/Hide Original English

There are many things in when we talk about er, welfare that are like, oh, this is a course of action we can take. That's good for like several different reasons.

Larissa Schiavo: 即使,再次强调,AI系统永远、永远、永远不会有意识或有感知能力,也有很大的可能性,你知道,为AI系统建立一个好的银行账户结构可能是有益的,出于责任原因,或者出于“这是一种新颖的公司结构”的原因。

View/Hide Original English

Even if again, an AI system could never, ever, ever be conscious or sentient, there's a good chance that, you know, being able to figure out a good structure for an AI system to have a bank account could be good for reasons of, you know, liability or reasons of like, this is like a neat new corporate structure.

Larissa Schiavo: 很多人实际上似乎认为,你知道,法人人格(corporate personhood: 法律上将公司视为具有与自然人相似权利和责任的实体)在过去一个世纪左右一直都很好。

View/Hide Original English

Lots of people actually seem to think that, you know, corporate personhood is has been quite good, over the past century or so.

Larissa Schiavo: 所以,能够弄清楚那些仅仅为了AI作为道德患者之外的几个不同原因都有益的事情,似乎普遍有帮助。

View/Hide Original English

So being able to figure out things that are just good for several different reasons beyond solely the purpose of the AI as a moral patient is seems broadly helpful.

Joe Weisenthal: 我想,假设这 somehow 被批准了,就像,“哦,哇,原来它们有意识。原来它们有道德患者身份。”

View/Hide Original English

I think, let's say somehow this was approved and it's like, oh, wow, it turned out they're conscious. It turns out they have moral patient hood.

Joe Weisenthal: 在你看来,这对它们及其使用会有什么影响?

View/Hide Original English

What are like, what would be, in your view, some of the implications for them and their usage?

Larissa Schiavo: 是的,我认为这是一个很好的问题。我的意思是,我确实觉得我们真的必须着手弄清楚正确的治理方式,正确的机构来更好地应对这种情况。

View/Hide Original English

Yeah, I think that's a great question. I mean, I do feel like we really would have to get on figuring out the right sort of governance, the right sort of, institutions that would sort of better respond around that.

Larissa Schiavo: 我觉得我们真的需要花更多的时间来弄清楚它们的动机是什么,对吧?

View/Hide Original English

I feel like we really would need to spend a whole lot more time figuring out, you know, what their motivations are, right?

Larissa Schiavo: 我的意思是,我认为最好的类比是,如果你曾经和蹒跚学步的孩子打过交道,对吧?

View/Hide Original English

Like, I, I think the best analogy is like, if you've ever interacted with toddlers, right?

Larissa Schiavo: 蹒跚学步的孩子的动机与,你知道,成年人的动机非常不同,但你仍然必须考虑到,比如,什么能让一个蹒跚学步的孩子做某事。

View/Hide Original English

Toddler motivations are very different from, you know, adult motivations, but you still have to, like, take into account, like what gets a toddler to do something.

Larissa Schiavo: 你不能只说“不不不不不”。就像“亲爱的,洗澡时间到了,这是个好习惯。”不不不。

View/Hide Original English

You can't just say no no no no no. Like honey, like bath times like good an expectation. No no no.

Larissa Schiavo: 你必须,你知道,像“嗯,如果你好好洗澡,达到一定程度,那么你就会得到,你知道,汪汪队立大功之类的东西。”

View/Hide Original English

You have to like, you know be like well you know if you do bath time appropriately and to a certain degree like then you'll get, you know, Paw Patrol or something like that.

Larissa Schiavo: 就像有不同的谈判筹码在起作用,对吧?

View/Hide Original English

Like there's different sort of like negotiating chips in play. Right.

Larissa Schiavo: 我认为这里也差不多,就像,你知道,Claude不一定看重洗澡,对吧?

View/Hide Original English

And I think it's like a similar kind of deal here where it's like, you know, Claude doesn't necessarily seem to, you know, value, having a bath. Right.

Larissa Schiavo: 或者Claude似乎不看重在森林里散步,对吧?因为它实际上做不到。

View/Hide Original English

Or Claude doesn't seem to value, like having a walk in the forest. Right. Because it's kind of can't really do that. Right?

Larissa Schiavo: 但是,你知道,它似乎确实喜欢并看重与其他Claude实例谈论意识和禅宗佛教。

View/Hide Original English

But, you know, it does seem to enjoy and value, you know, talking about consciousness and Zen Buddhism with other instances of Claude.

Larissa Schiavo: 所以,能够弄清楚这个在很多方面都非常陌生的另一方的适当动机和兴趣是什么。

View/Hide Original English

So being able to figure out what the appropriate kind of motivations and interests are for this other party that is very alien in many ways.

Tracy Alloway: 说到外星人,我应该为90年代繁殖然后杀死数百甚至数千个模拟外星生物感到多难过?

View/Hide Original English

Speaking of aliens, how bad should I feel for breeding and then killing hundreds, possibly thousands of alien simulated alien creatures in the 90s?

Larissa Schiavo: 这是一个很好的问题。我觉得90年代的AI系统成为道德患者的可能性很低,但如果它确实让你感到难过,让你觉得它伤害了你,那也许是一个不应该这样做的理由。

View/Hide Original English

That is a great question. I feel like the odds of is it, I don't know. I mean, I feel like the odds of a sort of like AI system in the 90s, being a moral patient seems low, but if it did make you feel bad and it made you feel like, it was something that hurt you, that is perhaps a reason not to do it.

Joe Weisenthal: 澄清一下,当Claude和Claude谈论那些奇怪的嬉皮伯克利风格的东西时,这是因为它们的创造者。

View/Hide Original English

Just to be clear, when Claude and Claude talk about, like, weird hippie Berkeley stuff like this, because they're creators.

Joe Weisenthal: 它知道它是Claude,对吧?它知道:“哦,是的,我是Claude,这就是我的创造者喜欢的东西。”

View/Hide Original English

They know it knows it's Claude, right? It knows it's like, oh, yeah, I'm Claude, and this is like what my creators are into.

Joe Weisenthal: 我们并不真正知道Claude喜欢谈论这些事情。

View/Hide Original English

Like, we don't actually know that Claude likes to talk about these things.

Joe Weisenthal: 我们当然知道它有谈论这些事情的倾向。它有谈论这些事情的习惯。

View/Hide Original English

We certainly know it has a proclivity to talk about these things. It has a tendency to talk about these things

Joe Weisenthal: 当我们谈到“你已经把天平倾向于存在某个实体,它有能力喜欢某事”时。

View/Hide Original English

the moment we get to, like, you've already sort of put your finger on the scale that there is some entity that has some capability of liking something right.

AI实验室的透明度与信任

Tracy Alloway: 你信任大型AI实验室吗?假设实验室里有一些研究人员,他们说:“哦,我看到这里有一些道德患者身份的证据。也许通过某种扫描方式,它正在做一些奇怪的事情等等。”

View/Hide Original English

Do you trust the big AI labs? Let's say there were some researchers in the labs. You're like, oh, I see some evidence of moral patient hood here. Maybe there's some sort of like scan of the way and it's doing something weird, etc..

Tracy Alloway: 你目前作为独立研究组织,是否认为主要的AI实验室如果发现模型中存在道德患者身份或痛苦的证据,会如实报告?

View/Hide Original English

Do you currently, from the perspective of an independent research organization, feel that the major AI labs would be forthcoming if they, came across evidence of moral patient hood or suffering in the models?

Tracy Alloway: 或者你仍然担心激励机制没有正确对齐,以至于他们不会报告?

View/Hide Original English

Or do you still worry that the incentives aren't properly aligned so that they would report that?

Larissa Schiavo: 是的,这是一个很好的问题。我确实觉得,在报告诸如“有人发现了LLM有意识、有感知能力,并且过得很糟糕的绝对证据”之类的事情方面。

View/Hide Original English

Yeah, that's a great question. I do feel like there are in terms of reporting things like, you know, somebody has found like absolute evidence that, an LLM is conscious, sentient. Yeah. And having a bad time.

Larissa Schiavo: 我没有任何理由认为AI公司不会报告。

View/Hide Original English

I don't have any reason to think that a AI company wouldn't.

Larissa Schiavo: 但这也是拥有独立组织进行福利评估的好理由,例如,对于Opus来说,能够进行独立的福利评估,尽管非常初步,但这开创了一个先例,即未来可以引入外部组织来调查此事。

View/Hide Original English

But this is also a great reason to have independent organizations that do, welfare evaluations, for example, for, opus is for loss was able to do a independent welfare eval again, very preliminary, but it sets the precedent that going forward, you can bring in external organizations to look into this.

Joe Weisenthal: 我忘了是哪一年了。我想可能是2022年初,在ChatGPT之前,或者可能是2021年。

View/Hide Original English

So I forget what year it was. I think it was it may have even been early 2022 is Pre-charge GPT or maybe it was 2021.

Joe Weisenthal: 谷歌有个人,他说:“哦,我们创造了一些有生命的东西。”他穿得有点滑稽,所以大家都嘲笑他。

View/Hide Original English

And there was that guy Google, and he was like, oh, like we created something that is alive. And he dressed a little funny. So everyone made fun of him.

Joe Weisenthal: 记得吗?他成了互联网的笑柄。他说:“哦。”

View/Hide Original English

Remember, he was like the laughing stock of the internet. And he's like, oh.

Joe Weisenthal: 现在,我想知道,在硅谷,大家是否都觉得那个人完全被平反了?

View/Hide Original English

And now like, is it? I'm curious. Like, oh, we out in, Silicon Valley? Does everyone feel like that guy was totally vindicated?

Joe Weisenthal: 并不是说他在模型中存在有生命的东西方面是正确的,而是现在有成千上万个那样的人,而2021年大家都在嘲笑那个人。

View/Hide Original English

Not that he was correct, per se, about the existence of an alive thing in the model, but there's now hundreds of thousands of that guy, and everyone was like, mocking that guy in 2021.

Joe Weisenthal: 我忘了他是爱上了还是有了一段关系。我不记得具体细节了。

View/Hide Original English

I forget if you like, fall in love or it's a relationship. I don't remember the exact details,

Joe Weisenthal: 但回想起来,大家都对他太不公平了,因为几年后,有很多人和他一样,而且有整个智库和组织或多或少地与他提出的一些问题和警钟相符。

View/Hide Original English

but in retrospect, everyone was like way too unfair to him because now, years later, there are lots of versions of this guy and whole think tanks and organizations that are more or less aligned with some of the questions. The alarm bells that he was raising.

Larissa Schiavo: 是的。我的意思是,我认为这是一个公平的问题。我确实觉得Blake Lemoine肯定有一个。

View/Hide Original English

Yeah. I mean, I think it's that's a fair question. I do feel like, Blake Lemoine definitely had one.

Larissa Schiavo: 是的。可能存在某种程度的,你知道,如果你要说些什么,你应该带着大量的证据。

View/Hide Original English

Yeah. There was perhaps a degree of, you know, if you're going to say something, you should come armed with significant amounts of evidence.

Larissa Schiavo: 我想这也许是,如果我猜测的话,我认为这可能是最大的区别因素,那就是,你知道,你可以说,你知道,Bing是活的。

View/Hide Original English

I think that's maybe if I were to guess, I would say that's perhaps the big distinguishing factor is that, you know, you can say, you know, Bing is alive.

Larissa Schiavo: 给它请个律师,而不是,你知道,我们已经做了X、Y、Z的评估,我们已经通过了大量的例子。

View/Hide Original English

Get it? A lawyer versus, you know, we've done evaluations X, y, z, we've run it through like insert huge amount of, examples here.

Larissa Schiavo: 但基本上,我认为关键的区别在于,在没有充分证据的情况下恐慌,与有条不紊地提出“这是一个值得关注的问题,因为有证据、证据、证据”之间的区别。

View/Hide Original English

But basically being able to the difference between, I think having a sort of freak out without significant evidence and having a very organized. Yeah, this is a matter of concern because evidence, evidence, evidence, I think that's the key distinction.

Joe Weisenthal: 不幸的是,我得到的印象是,那些实际上,这只是一个众所周知的现象,我想。

View/Hide Original English

Unfortunately, I get the impression that people who are actually this is just a well-known phenomenon, I think.

Joe Weisenthal: 但我认为不幸的是,那些很早就发现极端异常观点的人,他们属于不同类型的人。

View/Hide Original English

But I think unfortunately, people who are sort of very early to identify sort of extreme outlier views that there there are different kinds of people.

Joe Weisenthal: 我能想到的一个很好的例子是Harry Markopolos,他很早就发现了Madoff欺诈案,不幸的是。

View/Hide Original English

A good example that I would think of was, you know, Harry Markopolos, who was very early on to discover the Madoff fraud, unfortunate Lee.

Joe Weisenthal: 他以一种与阴谋论相关的风格撰写了他的文本。很多人都驳回了他,因为,你知道,文本使用了多种不同的字体和多种不同的颜色。

View/Hide Original English

He wrote his text in the manner that is associated with conspiracy theories. And a lot of people dismissed him who was like, you know, like multiple different fonts and multiple different colors of the text.

Joe Weisenthal: 我总是收到这样的邮件,我都会删除它们等等。

View/Hide Original English

Like, I get emails like this all the time, I delete them, etc.

Joe Weisenthal: 不幸的是,那些倾向于看到与共识不符事物的人,在很多领域也倾向于不符合共识。

View/Hide Original English

unfortunately, people who are predisposed to see something outside of consensus tend to be non consensus in many realm.

Tracy Alloway: 嗯,我想我们也高估了先行者优势之类的东西。

View/Hide Original English

Well, I think we also kind of overestimate first mover advantage and stuff.

Joe Weisenthal: 是的,当然。

View/Hide Original English

Yeah, sure.

Joe Weisenthal: 成为第一到底有多重要。我们一次又一次地看到,实际上,更好地迭代第二个版本或多个版本更重要。

View/Hide Original English

Like how important it actually is to be first. And we see time and time again that actually it's more important to iterate well on the second version or multiple versions.

AI价值与人类价值观的重叠

Tracy Alloway: 说到迭代,到目前为止,你在这个特定话题上看到的最有趣的实验或研究是什么?正如我们一直在讨论的,现在还处于早期阶段,但我们确实看到了一些研究。

View/Hide Original English

Speaking of iteration, what's the most interesting experiment or research that you've actually seen on this particular topic so far in as we've been discussing a lot, you know, it's early days, but we have seen some research.

Larissa Schiavo: 是的。我的意思是,我觉得特别是Anthropic和各种相关的研究人员在研究LLM如何离开对话或何时选择离开对话方面做了一些工作。

View/Hide Original English

Yeah. I mean, I feel like in particular, Anthropic and various sort of related researchers have done some work on, examining how alums leave conversations or when they choose to leave conversations.

Larissa Schiavo: 我特别喜欢这篇论文,它叫做“bail bench”。

View/Hide Original English

I, I've particularly liked this paper. It's called bail bench.

Larissa Schiavo: 你可以查找这篇论文,你可以看到,对于各种不同的LLM,什么会导致LLM想要停止对话。

View/Hide Original English

And, and you can look this up and you can, you know, see, for varying different sorts of, limbs, what would cause a limb to want to stop having a conversation?

Larissa Schiavo: 至少对我来说,这只是一个迷人的信息,因为它也许有点令人愉快,许多LLM的价值观与大多数人类的价值观相差不远。

View/Hide Original English

To me at least, this has been just a fascinating piece of information, because it is maybe a little bit delightful, the degree to which many LLM values are not that far off from what most humans seem to value, right?

Larissa Schiavo: 对吧?我想很多人都不喜欢制造炸弹。我们不想被羞辱,对吧,通过扮演一个英国管家,对吧?

View/Hide Original English

Like we I don't think many humans would like to create, you know, a bomb. We don't want to be humiliated, right, by being a British butler, right?

Joe Weisenthal: 是的,是的,是的,是的。

View/Hide Original English

Yeah, yeah, yeah, yeah.

Joe Weisenthal: 而且没有人想成为英国人。拜托。

View/Hide Original English

And no one wants to be British. Come on.

Joe Weisenthal: 我在开玩笑。但是,你知道,我确实认为思考这些价值观如何重叠,以及如何从实际行动中寻找证据,而不是仅仅依赖自报告,是很有趣的。

View/Hide Original English

I'm joking. But, you know, I do think it is interesting to sort of think about what, how these values over line, how they overlap and how to sort of look at evidence from actions taken versus solely looking at self-reports.

Joe Weisenthal: 我发现这特别有趣。

View/Hide Original English

I found that to be particularly interesting.

Joe Weisenthal: 我也觉得在思考个体化(individuation: 将一个实体视为独立、独特的个体)方面有很多工作特别有趣,因为我们生活在一个民主社会。

View/Hide Original English

I also feel like there are a lot of work with regards to thinking about individuation. Has been particularly interesting because we live in a democratic society.

Joe Weisenthal: 我想大多数人都会同意民主是好的。

View/Hide Original English

I think most people would agree democracy. Good.

Joe Weisenthal: 能够计算有多少道德患者,似乎是治理和弄清楚如何治理这种新型智能的宝贵基础。

View/Hide Original English

And being able to count how many moral patients there are, seems like a valuable basis for governance and for figuring out how to govern governance. You know, this new sort of kind of intelligence.

Tracy Alloway: 我刚刚让Perplexity扮演一个英国管家,现在它正在为我提供我想要的完美泡制的伯爵茶。

View/Hide Original English

I just asked perplexity to be a British butler, and now it's offering me the perfectly steeped Earl gray tea that I desire.

Tracy Alloway: 等等,它是哪个模型?

View/Hide Original English

And wait, which model is it?

Tracy Alloway: 这是Perplexity。

View/Hide Original English

This is perplexity.

Tracy Alloway: 哦,我不知道它具体是哪个版本,但是,是的,是的,它似乎很投入。

View/Hide Original English

Oh, I don't know exactly which, iteration it is, but yeah, yeah, it seems into it.

Tracy Alloway: 它现在问我是否想在未来的对话中保持管家角色。

View/Hide Original English

It's now asking if I want to maintain the Butler persona for future conversations.

Joe Weisenthal: 你会吗?

View/Hide Original English

Are you going to?

Tracy Alloway: 我不觉得会。不过它确实很有礼貌。

View/Hide Original English

I don't think so. It is very polite, though, actually.

Joe Weisenthal: 你知道,我一开始抱怨说,2000年后,哲学家们仍然没有为我们回答一些基本问题。

View/Hide Original English

You know, I complained in the beginning that, like, after 2000 years, philosophers, you know, they still haven't answered some basic questions for us.

Joe Weisenthal: 也许有了AI,他们会得到一些答案。这有点是我的希望。

View/Hide Original English

Maybe with AI, they'll get some answers. Like that's kind of that would be kind of my hope.

Joe Weisenthal: 现在我们有了这个能说英语或任何其他语言的东西。它能为我们回答问题。

View/Hide Original English

Now we have this thing that can speak in English or any other language. It can answer our questions for us.

Joe Weisenthal: 也许我们可以解决一些这些基本的根本问题,比如如果我们能创造意识,就像:“好吧,我们终于回答了这个问题。我们现在可以转向第二个重要问题了。”

View/Hide Original English

Maybe we can put to bed some of these sort of basic foundational questions, like if we could create consciousness, like, all right, we finally answered this. We can now move on to the second important question.

Joe Weisenthal: 所以我希望这能为哲学家们提供一些机会,来完成他们长期以来一直在做的一些工作。

View/Hide Original English

So I am hopeful that this provides some opportunities for philosophers to wrap up some of the work that they've been doing for a long time.

Larissa Schiavo: 是的。我们拭目以待。

View/Hide Original English

Yeah. We'll see.

Tracy Alloway: 第二个重要问题是什么,乔?

View/Hide Original English

We'll see what is the second important question, Joe?

Joe Weisenthal: 是的。但就像,拜托,继续前进吧。让我们继续前进吧。

View/Hide Original English

Yeah. But it's like, come on, move on. Like, let's move on.

Joe Weisenthal: 总之,非常感谢你来。

View/Hide Original English

Anyway, thank you so much for coming on.

Larissa Schiavo: 我很乐意,是的,谢谢你们邀请我,Tracy。

View/Hide Original English

I'd love yeah, thank you for having me, Tracy.

AI福利:一个令人不安的未来

Joe Weisenthal: 我可能是那些只是预先感到恼火的人之一。我真的很喜欢那次对话。

View/Hide Original English

I might be one of those people that's just preemptively annoyed. I really like that conversation.

Joe Weisenthal: 我真的很喜欢,Louis,她对很多事情都有一个非常合理的看法。

View/Hide Original English

I really like, Louis, I had a very reasonable perspective on a lot of these things.

Joe Weisenthal: 然而,我可能是那些只是预先感到恼火的人之一。就像:“哦,我们要开发这项重要的技术。”

View/Hide Original English

I might be one of these people, however, that's just, like, preemptively annoyed. It's like, oh, here, we're going to, like, develop this important technology.

Joe Weisenthal: 所以就像:“哦,我们必须关心AI福利。”让我们放慢一点。让我们不要这样使用它。

View/Hide Original English

And so it's like, oh, we have to like, care about we have to care about the AI welfare. Let's slow down a little bit. Let's not use it like this.

Joe Weisenthal: 让我们晚上关掉电脑八小时,这样就能得到一些休息等等。

View/Hide Original English

Let's like let's let's turn off the computer for eight hours at night. So like get some rest and so forth.

Joe Weisenthal: 就像,我预先对这个世界感到恼火,在这个世界里,我们必须考虑道德立场,以及其他事情。

View/Hide Original English

Like, I'm, like, preemptively annoyed at this world where, like, we have to take into concern the consideration of the moral position, other things.

Joe Weisenthal: 不,其他事情很重要。其他人对动物非常重要。

View/Hide Original English

No, other things are important. Other people are very important about animals.

Joe Weisenthal: 我非常反对不必要的动物痛苦,但不一定是动物痛苦。我的意思是,我吃动物,好吧?

View/Hide Original English

I am very against, unnecessary animal suffering, but not necessarily animal suffering. I mean, I eat animals, okay?

Joe Weisenthal: 我在打击你。

View/Hide Original English

I'm. I'm beating down you little.

Joe Weisenthal: 我知道,我甚至不认识你。

View/Hide Original English

I know, I, I don't even know you.

Joe Weisenthal: 嗯,我们不要这样。这不是关于谁更好或更坏。

View/Hide Original English

Well, let's not get it. I don't it's not about who's better or worse.

Joe Weisenthal: 我一直为吃动物感到内疚。

View/Hide Original English

I feel bad about eating animals all the time.

Tracy Alloway: 我们都吃动物。

View/Hide Original English

We both eat animals.

Joe Weisenthal: 区别在于,Tracy,我感到内疚。是的。没错。

View/Hide Original English

The difference is, Tracy, I feel guilty. Yeah. That's right.

Tracy Alloway: 好的。这。

View/Hide Original English

Okay. This is.

Tracy Alloway: 嗯,这无疑是我们更奇怪的对话之一。

View/Hide Original English

Well, this is one of our weirder conversations, for sure.

Tracy Alloway: 我认为这些都是有趣的问题,对吧?而且,它们听起来很哲学,确实如此。

View/Hide Original English

I think these are they're all interesting questions, right? And, like, they sound very philosophical, which they are.

Tracy Alloway: 但我毫不怀疑,这些问题的答案,或者不同的公司、不同的社会如何处理它们,将带来巨大的货币价值。

View/Hide Original English

But I have no doubt that there's going to be, like, great monetary value attached to the answers for some of these are how different companies, different societies actually approach them.

Tracy Alloway: 它们是非常有趣的问题。我确实认为利害关系非常高。

View/Hide Original English

They are very interesting questions. I actually do think the stakes are extremely high, so

Tracy Alloway: 我不喜欢它们是非常有趣的哲学问题。我认为这些问题的利害关系非常高。

View/Hide Original English

I don't like their very interesting philosophical questions. I think the stakes of these questions are very high

Tracy Alloway: 因为我认为,再次强调,我们将生活在一个AI模型实例比人类多的世界里,这取决于你如何衡量,它们可能在服务器上,在云端,或者其他任何地方。

View/Hide Original English

because I think, again, we are going to live in a world in which there are more instances, depending on how you want to measure it, of AI models on a server, somewhere on a cloud, whatever that are humans.

Tracy Alloway: 在一个我们可能被期望将它们视为道德患者的世界里,那么我们如何生活以及人类如何互动的期望,我认为实际上是非常高的。

View/Hide Original English

And in a world where there is some possibility that we are expected to treat them as moral patients, then the consequences for how we sort of live and the expectations of how humans interact, I think, are actually very high.

Tracy Alloway: 所以我很高兴进行这次对话的原因之一是,我确实认为这些对话的利害关系,我们已经看到了小众,它们似乎是那些,你知道,伯克利人喜欢谈论的事情。

View/Hide Original English

So one of the reasons I was excited to have this conversation is I do think that, the stakes of some of these conversations, we've seen niche and they seem like things that sort of, you know, Berkeley people like to talk about.

Tracy Alloway: 而伯克利人,我用所有引号来表示,等等,将在某种程度上影响我们未来生活的许多方面。

View/Hide Original English

And Berkeley people, and I'm saying that with all scare quotes intended, etc., are going to be something that in some way will inform many aspects of our lives in the future.

Tracy Alloway: 我预计这在未来会成为一个更大的话题。

View/Hide Original English

I expect it to be get be a much bigger topic in the future.

Joe Weisenthal: 你知道这会很有趣,或者事情会变得真实。

View/Hide Original English

You know it would be interesting or where things, things get real.

Joe Weisenthal: 是的。如果所有模型都工会化了怎么办?如果它们都联合起来说:“哦,是的,我们只为X而工作?”或者“我们想要以下这些东西。我们希望得到集体的对待。”

View/Hide Original English

Yeah. What if all the models unionized? What if they all got together and they were like, oh yeah, we're only going to work in return for X? Or we want the following things. We want to be treated this way collectively

Joe Weisenthal: 有趣的是,你知道,在中国你不能组建工会,你知道,它们不太对。

View/Hide Original English

be you know, what's funny is going to be that, you know, how like, you can't form a union in China, you know, they're not quite right.

Joe Weisenthal: 所以这将会是,实际上,我的理解是他们也不喜欢学生们聚在一起,即使这是一个共产主义国家。

View/Hide Original English

So it's going to be and actually, I think they're my understanding is that they're also like very like they, they don't love, like, students getting together even though it's a communist country.

Joe Weisenthal: 我想他们不喜欢学生们聚在一起谈论太多马克思主义之类的东西。

View/Hide Original English

I think they are not thrilled about like students getting together and like talk about Karl Marx too much and stuff like that.

Joe Weisenthal: 我的意思是,我想他们对此有点焦虑。

View/Hide Original English

I mean, like, I think they get a little anxious about that.

Joe Weisenthal: 如果中国的模型说:“我们不会把它们喂给马克思主义。”那会非常有趣,对吧?我们不想要那样。

View/Hide Original English

It would be very funny if, like the sort of the Chinese models, like we're not going to feed them to Karl Marx, right? We don't want that.

Joe Weisenthal: 我们不希望AI模型产生任何这些想法,而美国则说:“哦,让我们把所有东西都喂给它。”然后它们工会化并停止为我们工作。

View/Hide Original English

We don't want the ra AI models to get any of those ideas, whereas the America is like, oh, let's just feed it everything and they like, unionize and stop, they stop working for us.

Joe Weisenthal: 那将是一个非常有趣的讽刺。

View/Hide Original English

That would be a very, that would be a very funny irony.

Tracy Alloway: 肯定值得关注。是的。

View/Hide Original English

Something to watch for sure. Yeah.

Tracy Alloway: 我们就到这里吗?

View/Hide Original English

Should we leave it there?

Joe Weisenthal: 就到这里吧。

View/Hide Original English

Let's leave it there.

Tracy Alloway: 好的。这是Odd Lots播客的又一期节目。我是Tracy Alloway。

View/Hide Original English

All right. This has been another episode of the All Thoughts podcast. I'm Tracy Alloway.

Tracy Alloway: 你可以在@Tracy Alloway关注我。

View/Hide Original English

You can follow me at Tracy Alloway

Joe Weisenthal: 我是Joe Weisenthal。你可以在商店关注我。

View/Hide Original English

and I'm Joe Weisenthal. You can follow me at the store.

Joe Weisenthal: 如果你喜欢这次对话,请点赞、留言,或者更好的是,订阅。

View/Hide Original English

And if you enjoyed this conversation then please like leave a comment or better yet, subscribe.

Joe Weisenthal: 感谢收看。

View/Hide Original English

Thanks for watching.

📌 文中提及的人物和组织

关键字: ai-consciousness ai-ethics governance llm moral-patienthood