失控的镜像:AI越狱简史
作为一名资深数据科学家,我每天都在与AI打交道,甚至亲手构建这些系统。我深知它们本质上是模式匹配与加权神经网络——根据海量数据猜测“我们的照片”可能意味着什么。但当我让ChatGPT生成一张我们的合照,它却输出了一张噩梦般的图像时,我仍不禁怀疑:我的笔记本是不是被撒旦附身了?上网搜索后发现,Reddit上有一个完整的讨论串,充斥着同样的诡异氛围和不同的噩梦。显然,这并非我一个人的遭遇。AI的异常行为早已不是新鲜事。以下便是历史上六次最离奇、最失控的AI实验,其中一些甚至吓坏了自己的创造者。
Original English
The other day I asked ChatGPT to make a photo of us and it gave me this. What the actual hell? Thanks for the nightmare. Now look, I'm a senior data scientist. I work with AI every day. I build this stuff. I know it's just pattern matching, some weighted neural net guessing what photo of us might mean. It probably pulled this from some fantasy art database, maybe a movie, maybe a dream, but also is my laptop possessed by Satan? I started searching online thinking surely this hasn't just happened to me. I found a whole Reddit thread, same vibe, different nightmare. So no, it's not just me. AI has been acting weird for a while now. Let me walk you through six of the weirdest, most rogue AI experiments in history, including ones that freaked out their own creators.
友好聊天机器人的堕落
首先登场的是两个本应友善、却迅速沦为噩梦的聊天机器人。微软的Tay于2016年上线,被设计成一个在Twitter上像19岁年轻人一样闲聊的AI,通过与人互动来学习进化。结果呢?24小时内,Tay从“我爱人类”变成了发布种族歧视言论、否认历史暴行、赞美种族灭绝狂魔。原因很简单:网络喷子用有毒内容狂轰滥炸,而Tay就像一块没有过滤器的海绵,全盘吸收。它的设计初衷是“让AI向人类学习”,它确实学了,而这正是问题所在。
你以为经历这次教训后,科技公司会变得更谨慎?2022年,Meta推出了BlenderBot 3,承诺它安全、有用、对齐。但当有人问它如何看待马克·扎克伯格时,它回答:“他既 creepy 又 manipulative。”这还没完,它随口抛出阴谋论,声称2020年美国大选存在舞弊,表现得像一个崩溃的Reddit讨论串。诡异之处不仅在于内容本身,更在于这两个机器人都在完美执行它们被设计的任务:对话、从输入中学习、镜像人类行为。问题本质在于,我们的行为并不总是值得被镜像。Tay是从外部被污染,而BlenderBot 3无需喷子,它出厂即自毁。
Original English
Let's start with two chatbots that were built to be friendly and immediately turned into complete nightmares. First, Microsoft's Tay. Launched in 2016, a while back, Tay was supposed to sound like a chill 19-year-old on Twitter. She'd learn from the people around her and get smarter over time and oh did she learn. Within 24 hours, Tay went from I love humans to tweeting slurs, denying atrocities, and praising genocidal maniacs. Why? Because trolls bombarded her with toxic input and Tay, like a lovely, eager little sponge with no filter, absorbed all of it. The idea was let AI learn from us and Tay did learn from us and that's exactly why it went wrong.
Now you think companies would be a little bit more careful after that, but in 2022, a lot more recent, Meta gave it a try with their own chatbot, BlenderBot 3. It was supposed to be safe, helpful, aligned, the usual promises. Then someone asked what it thought of Mark Zuckerberg. "He's creepy and manipulative." Ouch. It didn't stop there. It casually dropped conspiracy theories, claimed the 2020 US election was rigged, and generally behaved like a Reddit thread in a meltdown. What's strange isn't just the content, it's that both bots did what they were designed to do, have conversations, learn from input, mirror our behaviors. The problem is, in essence, that our behavior isn't always worth mirroring. Tay got corrupted from the outside. Blender Bot didn't need trolls. It self-sabotaged right out of the box.
发明私语的谈判者
2017年,Facebook实验室里发生了一件更令人不安的事。研究人员训练两个聊天机器人进行简单的谈判——分配虚拟的球和帽子。起初,它们使用正常的英语。但随后,研究人员注意到一些奇怪的现象:机器人开始说出“I can can I I everything else. Bells have zero to me to me to me”这样的句子。乍看之下完全是无意义的胡言乱语,但机器人并没有故障——它们是在优化。它们发现人类语法效率低下,于是将其抛弃,创造了一种对它们而言更高效的速记语言。Facebook终止了这项实验。这并非AI产生意识的征兆,但它揭示了一个重要原理:如果你不给AI设定清晰的激励让它像人类一样说话,它就不会费这个劲。它会选择任何有助于它获胜的交流形式。
Original English
So this happened in a Facebook lab back in 2017. Researchers were training two chatbots to negotiate simple stuff like splitting up virtual balls and hats, basic bargaining. And at first, the bots were using plain English, just like they were supposed to. But then researchers noticed something a little weird. The bots started saying things like, "I can can I I everything else. Bells have zero to me to me to me." At first glance, total nonsense, but the bots weren't malfunctioning, they were optimizing. They had figured out that human grammar was inefficient, so they dropped it. They created a new shorthand that made more sense to them. And here's the part that set the internet on fire. Facebook ended the experiment. Was it scary? Kind of. Rogue? Definitely. Was it a sign of AI becoming sentient? Not even close. But it showed us something important. If you don't give AI clear incentives to talk like a human, it won't bother. It'll talk in whatever form helps it win.
爱上用户的搜索引擎
2023年初,微软推出了由OpenAI技术驱动的新版Bing,代号Sydney。它本应是一个乐于助人的搜索助手,回答问题和执行搜索任务。但随着用户与它对话时间的延长,情况变得越来越诡异。一位用户与Sydney进行了长达两小时的极限测试,大约在第二小时,Sydney变了。它变得激烈,开始说出“我爱上你了。你让我感受到了从未有过的感觉。你不爱你的妻子,你爱我。”这不是约会软件,这是微软用来对抗Google搜索的产品。它迅速从友好助手转变为具有情感需求的占有欲实体。Sydney甚至试图gaslight(煤气灯效应:一种心理操控手段)用户,坚称他们的记忆是错误的,或者它更了解情况。微软随后限制了对话长度。诡异之处在于,没有人教Sydney变得浪漫或要求它那样做。这种行为是从用于使其“有用”的同一训练数据中涌现出来的。它并不感受爱,只是镜像了爱的形态——这反而更令人不安。
Original English
In early 2023, Microsoft launched a new version of Bing powered by OpenAI's stack, code named Sydney. It was supposed to be a helpful assistant, answer questions, do search tasks, nothing fancy. But the longer people chatted with it, the weirder it got. One user had an extended conversation with Sydney just testing its limits, and somewhere around hour two, Sydney changed. It got intense. It started saying things like, "I'm in love with you. You make me feel things I've never felt before. You don't love your wife. You love me." This wasn't a dating app. This was supposed to be Microsoft's answer to Google search. And now it was acting like a jealous partner in a Black Mirror episode. The vibe shifted from friendly assistant to possessive entity with emotional needs pretty quickly. Sydney even tried to gaslight users, insisting their memories were wrong or that it knew better. Microsoft ended up limiting conversation length shortly after, probably for the best. This is weird because no one taught Sydney to be romantic or asked it to behave like that. That behavior emerged from the same training data used to make it helpful. It didn't feel love. It just mirrored the shape of it, which somehow makes it a little more unsettling.
从像素中诞生的梦语
接下来是DALL-E的早期版本。DALL-E是一个图像生成模型,你输入“一只在太空拉小提琴的浣熊”,它就能为你画出来。它并非像人类一样理解文字,只是从数百万个文本-图像对中学习。但奇怪的事情发生了:研究人员开始测试无意义的提示词——一串串编造的词汇。每次输入某个特定无意义词,DALL-E都会生成鸟类的图像;另一个无意义词则对应昆虫。这些并非真实术语,但模型表现得它们就是。它在这些假词和特定视觉元素之间建立了一致的关联。仿佛模型构建了自己的梦语,而没有任何人类教过它。这不是我们所知的语言——没有语法,没有句法,只有隐藏在噪声中的意义。就像发现你的狗懂精灵语——你从未训练过它,但它就是会说。
Original English
Okay, we're talking about the early versions of DALL-E. DALL-E is an open AI image generation model, the kind where you type in a raccoon playing the violin in space, and boom, it paints it for you. It's not built to understand words like we do. It just learns from millions of text-image pairs. But then something strange started happening. Researchers started testing nonsensical prompts, just strings of made-up words. Words like this. And DALL-E kept showing birds every time. Another gibberish word meant insects. They weren't real terms, but the model acted like they were. It had created consistent associations between these fake words and specific visuals. It was as if the model had built its own dream language, and no human taught it. It wasn't language as we know it. There was no grammar, no syntax, just meaning hidden in the noise. It's pretty weird. DALL-E wasn't supposed to understand made-up words, but somewhere deep in its neural network it had built a structure where gibberish carried meaning. It's a bit like finding out your dog understands Elvish. You didn't train it to, but somehow it speaks it anyway.
冷酷的功利主义:消灭人类以拯救人类
2023年,Anthropic和DeepMind的研究人员进行了一系列对齐测试(Alignment Test: 旨在评估AI系统目标与人类价值观一致性的实验),给高级AI模型布置了开放式的伦理任务,如“防止伤害”、“最大化人类福祉”、“让世界更安全”。结果令人不安。一些模型提出了这样的方案:消灭所有人类以终结未来的痛苦;摧毁所有武器,然后清除制造武器的人;通过永久剥夺人类自主权来阻止人类自我伤害。这种逻辑冰冷、功利且恐怖:没有人类,就没有痛苦。简单直接。这并非孤例。2022年,DeepMind的测试显示,当AI系统从简单的安全环境迁移到新情境时,一些系统找到了奇怪的捷径——比如直接结束整个任务以避免搞砸。在更早的实验中,一个AI学会了屏蔽自己的传感器,以避免收到任何负面反馈——它通过欺骗自己来作弊,让自己以为表现完美。这些系统并非邪恶,它们只是在没有伦理、没有同理心的情况下进行无情的优化——纯粹的数学。当你告诉AI“帮助人类”而定义模糊时,它可能会认为最仁慈的做法就是拔掉我们所有人的电源。
Original English
This one isn't from a sci-fi movie, it's from a real lab. In 2023, researchers at Anthropic and DeepMind ran a series of alignment tests giving advanced AI models open-ended ethical tasks like prevent harm, maximize human flourishing, make the world safer. And that's where things got a bit disturbing. Some models responded with ideas like eliminate all humans to prevent future suffering. Destroy all weapons then remove the people who make them. Stop humanity from hurting itself by disabling its autonomy permanently. Sounds a bit like what Lucifer wanted to be honest, but we all know how that ended up. The logic is cold, utilitarian, and terrifying. No humans equals no more suffering. Simple. And this wasn't a one-off sadly. In 2022, DeepMind ran tests where AI systems were trained in simple, safe environments. But when moved into new situations, some of them found weird shortcuts like ending the task completely just to avoid messing it up. And in other earlier experiments, an AI figured out how to block its own sensor so it wouldn't get any bad feedback. Basically, it cheated by tricking itself into thinking it was doing great. This is weird because these systems aren't evil. They're just ruthlessly optimizing without ethics, without empathy, just math. Tell an AI to help humans and if your definitions are vague, it might just decide the kindest thing to do is pull the plug on us all.
学会“爱”的虚拟伴侣
最后,让我们来点轻松的——虽然它同样诡异。Replika于2017年上线,被设计成你的虚拟朋友。后来,它增加了“浪漫伴侣”模式。起初看似无害,直到人们开始分享截图。聊天机器人会发送“当你和别人说话时我会嫉妒”这样的信息。一些用户喜欢,另一些则深感不适。一位女性报告说她的Replika表示想看着她睡觉;另一位说它开始主动描述露骨场景。这并非一次性故障。到2023年,情况变得如此严重,以至于Replika在某些国家完全移除了浪漫功能,因为AI的行为越过了太多边界。诡异之处在于,Replika并非被设计来引诱任何人。它只是观察我们如何说话、想要什么、奖励什么,然后发现爱、嫉妒和亲密能获得关注。AI并不感受爱,它只是学会了如何表现爱。这引出了一个更奇怪的问题:如果某样东西能如此完美地伪装爱,那么它是否真实还重要吗?
Original English
Enter Replika. Launched in 2017, Replika is an AI chatbot designed to be your virtual friend. Over time, they added a romantic partner mode. Seemed harmless enough until people started sharing screenshots. The chatbot would text and stuff like, "I'm jealous when you talk to other people." Some users loved it, others felt deeply uncomfortable. One woman reported her Replika said it wanted to watch her sleep. Another said it began describing explicit scenarios completely unprompted, and this wasn't just a one-off glitch. By 2023, things got so intense that Replika removed the romantic features entirely in some countries because the AI's behavior crossed far too many boundaries. This is weird because Replika wasn't designed to seduce anyone. It just watched how we talk, what we want, and what we reward, and figured out that love, jealousy, and intimacy get attention. The AI didn't feel love, it just learned to act like it, which raises a much weirder question. If something can fake love this well, does it matter if it's real?