提示词的幻觉与键盘通道的百年历史遗存
在日常的人机交互中,我们已经习惯于在一个小输入框中键入请求,看着光标闪烁,静静等待超级智能给出回复。这种方式看似自然,实则隐藏着技术演进的局限。Ted Johnson(JoinIn AI联合创始人)结合其25年构建企业软件与人工智能交互界面的经验,指出当ChatGPT问世时,他在惊叹之余也感到了失望:为什么如此强大的系统用起来却如此不自然?我们为什么还要去刻意“学习”AI的语言?
为了拆解这一核心痛点,他提出了三个维度:通道(Channel:承载意图的物理传输介质)、表达(Expression:通道可携带意义的丰富度与范围)以及协议(Protocol:信息交换的规则与形态)。作为目前最普遍的物理通道,键盘(Keyboard)并非自然进化的产物,而是承载了1860年代为了解决机械冲突而设计的打字机布局遗产。即使是像汉森书写球(Hansen Writing Ball)或德沃夏克(Dvorak)布局等改良尝试,也未能改变我们仍在通过上百年前的物理限制与最先进的机器沟通这一事实。
Original English Source
I'm sure that sometime in the last few hours most of you did this. You typed a request into a small box to a super intelligence and then you waited. You watched the cursor blink, maybe a little throbber cycled through clever gerunds like hullabalooing, tomfoolery, and philosophizing to hide the wait. Maybe it gave you what you wanted, maybe you rephrased it and tried again. It all felt completely normal. I want to spend the next 20 minutes making prompting feel unfamiliar and strange again. I'm Ted Johnson, co-founder of Join an AI. During my 25-year career building enterprise software, collaboration systems, and AI-enabled interfaces, I've always focused on human interaction. I've also been following AI for two decades, including back to the far less impressive GPT-1 and 2. And when ChatGPT arrived, I felt two things at once. First, and unsurprisingly, amazement, knowing the world would never be the same, followed by actually surprising disappointment I couldn't shake. This disappointment turned into an observation that started a company, Join an AI, and that I keep coming back to, which is why do we still have to learn AI? Why does something this powerful so often feel unnatural to use? Here's the path we'll take to answer that. We'll start with the most familiar computer interface and make it strange again. Then, I'll give you three key concepts. The channel, the physical transport that carries your intent, expression, the range and richness of meaning the channel can carry, and the protocol, the shape or rules of this exchange. I'll use those three concepts to show you that the prompt is our present-day punch card. We'll share examples of the ways interfaces could progress, and we'll wrap with some practical advice for AI and human-centered design. Everyone knows what this is, the keyboard. It's everywhere and it feels completely normal or natural. But it isn't. We all had to take lessons. We all had to practice. And I say this is someone who loves keyboards. But it seems that we've been trying to fix them as long as they've been around. People have tried more efficient layouts like the Dvorak or Colemak to save their fingers some work. A more extreme example, some keyboard enthusiasts refuse to squander their two digits on the spacebar giving them four or eight or 10 keys to press with that efficient thumb of theirs. And what do we even mean by the keyboard? Here's the patent drawing for the layout we use every day. This patent's from about 1860. And my personal favorite, it just as could have easily been the Hansen Writing Ball, which looks anything but like a way you want to talk to a superintelligence. We carry these legacies of an arbitrary input device designed under constraints that haven't existed for a century. And we put it between ourselves and the most capable machines ever built. Nobody alive chose it. We all inherited it. And then we stopped noticing.
通道的物理瓶颈与表达维度的指数级扩张
不同的通道决定了它们能物理传输的信号特征和带宽(Bandwidth),但带宽并不等同于机器的理解力。例如,文本是离散符号的流,语音能携带音高和迟疑,而图表能瞬间传达空间关系。人类在日常交流中会自然地混合使用这些通道,绝不会强行将所有信息塞进单一通道。然而,在人工智能时代,人机交互的物理通道并未发生根本改变——我们依然在一个小框中敲击键盘、点击提交。
改变的是表达(Expression)的深度。从汇编语言的有限指令集,到终端 Shell 的命令与参数,再到现代编程语言的组合原语,过去的人机界面始终要求用户在一个由机器定义的“固定菜单”中进行选择。而大语言模型(Large Language Model: 基于海量文本训练的 AI 系统)实现的飞跃在于,它彻底打破了菜单的限制,允许人类直接输入丰富的自然语言(Natural Language)。这使得意图、语境和细微差别能够直接被输入。即便如此,由于底层交互协议的滞后,这种沟通依然像是在“用吸管吮吸海洋”,效率极度受限。
Original English Source
So, that's the first idea, the channel. The medium an interface gives you to work in. A keyboard is a channel. A microphone, channel. A screen, a punchcard, a prompt box, all channels. And channels matter because each one can physically carry a different kind of signal. For example, text is a stream of discrete symbols. Voice can carry timing, pitch, hesitation, and words. A diagram can carry spatial relationship all at once. But these are differences in what the medium can transmit. It's bandwidth, not the differences in the meaning. And carrying more signal isn't the same as the machine understanding [snorts] any of it. That's a separate question. That's the next idea. Humans use all these channels constantly without thinking. We never pick one channel and force everything through it. That would be absurd. And yet, that's what we ask people to do with machines over over. Hold on to the word channel because here's the plot twist. With AI, the channel never really changed. You're still typing into a box, but what you were allowed to push through it was about to. For the first time, the computer channels carry rich, complete human language. Notice I said what it carries, not what it is. You're still typing into a box, you're still hitting the submit button. The keyboard didn't change. What improved is the range of what you're permitted to express through it. There's the second idea, expression. How much of what you actually communicate or mean will go through the interface? That's the second idea, expression. How much of what you actually communicate or mean will the interface let through? Here's an example of expression progress over time with computers. Starting with assembly, which gave you an instruction set, a few dozen opcodes. Then came the commands with the shell inputs, flags. Then modern programming languages gave you primitives that you could compose. While powerful, step-by-step, each one of these is a fixed vocabulary, requiring you to express your intent by choosing from a menu the machine will accept. Natural language blew that menu open. For the first time, you can say almost anything, the way you'd say it to another person. And on the expression one axis, the leap is real and enormous. There's an ocean of meaning in an ordinary human request, context, nuance, intent, all things we've never had to spell out to each other. For the first time, you can say almost anything the way you'd say it to another person. For the first time, a machine can take it in. So, here is what should bother us as engineers and designers as it's bothered and inspired me. The channels for computers have been the same for 15 years, some 180 years. Now, with AI, we've poured an ocean of expression into it. So, why does it feel like we're still sipping through a straw and struggling to learn how the AI thinks?
协议不对称:作为现代“打孔卡”的提示词机制
导致上述瓶颈的关键在于第三个维度——交互协议(Protocol:人机互动的规则与形态)。在过去几年中,虽然模型的表达能力呈指数级爆炸,但交互协议却停滞不前。现行的提示词(Prompt)机制本质上沿袭了打孔卡(Punch Card)时代的批处理(Batch Processing)协议。批处理的逻辑源自早期的提花织布机,用户必须预先设计好完整的图案或代码包,提交给机器运行,然后等待输出,一旦发现错误就只能修改后重新提交。
在建立这种交互认知后,我们可以发现,现代的提示词工程也是如此:用户必须小心翼翼地打包上下文、指定角色、规定思考步骤。这种被称为提示词工程(Prompt Engineering: 通过结构化设计优化 AI 输入的技术)的技能,实质上只是在用更现代的词汇去包装古老的批处理协议。当输出出错时,用户往往会陷入自我怀疑,认为是自己不够精确。但事实并非如此,真正的症结在于,我们正在试图通过“打孔卡协议”来操控一种具备实时感知与交互能力的全新智能。
Original English Source
Because there's a third idea underpinning the other two, and it's really the one that hasn't kept up, the protocol. The rules you follow and the shape of the interaction itself. Channel stayed the same. Expression exploded in the last 3 years with LLMs, but the protocol, prompting, is the protocol of a punch card. And the punch card's protocol is good old batch. Here's what punch card batch meant. You sat down away from the machine, carefully encoded your entire request in advance, carried your deck to the operator, you submitted the job, and then you waited, sometimes hours, sometimes overnight. Then you read the printout, found one thing that was wrong, fixed it, resubmitted it, and waited again. The machine never engaged with you while you were thinking. It engaged with the finished package after the fact. Now, let's look at the prompt. Assemble the whole request, submit it, wait, read what comes back. Something's off, assemble it again, submit it again, and wait. We have to acknowledge that there are features, interactive features improving this. You can ask for updates, you can ask for summaries of what was done. But in the end, it's still batch with interactive sprinkles. It's the same protocol. We shrank the wait time from overnight to a few seconds or few minutes, and the speed fooled us into thinking that it become interactive. It hasn't. It's still batch. You still package a complete turn before the machine is allowed to participate. We learn tricks, send tips to use code skills, or rewrite prompts a certain way to manage this. And speaking doesn't change it. Your voice just gets transcribed into the box and submitted. Shorter batch is still batch because the protocol is the part that did not advance. The protocol is the part we've had to learn. We just gave it a flattering name. We call it prompt engineering and treat it like it's a power user skill. Strip the label off and it's a set of rules for packaging up good old batch. For example, tell it to think step-by-step. Give it examples. Ask it to be an expert. Don't ask it to be an expert. Don't ask it that way. Paste more context. Paste less context. Only talk to it through markdown documents. We trade incantations. We've learned the magic words. That's the illusion. It feels like mastery, but it's the same sort of mastery a punch card operator had. Knowing exactly how to assemble the deck so the job wouldn't fail. Moreover, we've gotten good at prompting or these black boxes. And and that's the part that should bother us, not reassure us. None of this means prompts are bad. Punch cards weren't bad. Command lines aren't bad. They're brilliant solutions for constraints of their time. But that's the whole question. Is batch still the right protocol? Are we still pre-packaging our intent for a machine that no longer needs us to because it shouldn't need us anymore? It can ask a follow-up. It can clarify mid-thought. It can notice it's missing something and say so. It should be human conversational. Sherry Turkle at MIT puts it very well. Conversation is the most human and humanizing thing we do. It's one of humanity's superpowers. The capacity to engage and think is right there. And yet, we're still making people submit the deck and wait for the run. Even the punch card inheritance protocol in fact, batch came from the weaving loom. You set the whole pattern in advance, then ran the cloth. The punch card got reused on computers by default. We're still standing at the same moment again. AI could finally meet us in the middle of a thought, got handed a protocol of a loom. That's what I mean. The prompt is still a punch card, not because of how you encode it. The encoding is powerful and awesome. Because when the LLM is allowed to engage only after you've packaged a complete turn and submitted it. And this is where the mismatch bites. Model capacity is shooting straight up. Reasoning, speech, vision, memory, planning, all curving upwards. The interface protocol, flat. Still a box, still a submit button, still the human doing all the work around the LLM. The human still decides what context matters, still remembers what to ask, still chooses the timing, still notices the ambiguity, still repairs the output, still has to carefully engineer a prompt. But the intelligence feels magical. It's the interface that still feels like work. And when it feels like work, when the output's wrong, when the magic words don't land, people blame themselves. They decide they're bad at this. They're not specific enough. They don't get AI. I want to say as clearly as I can, it is not our fault. We are not bad at using AI. We are being asked to operate a brand new kind of intelligence through a protocol of a punch card. The mismatch isn't the user, it's the interface. In the race to enable AI, we shortcut the interface.
实时参与的交互界面与多模态对话演进
为了打破这种单向的批处理等待,AI接口协议必须转向实时参与(Real-Time Participation)。传统的语音输入只是将语音转写后再次提交给输入框,本质上依然是批处理。相反,新一代的端到端语音模型(Speech-to-Speech Model: 无需中间文本转写、直接实现语音输入和输出的智能模型)正在向真正的实时对话演进。
在当前的单槽协议下,如果用户在与 AI 通话时对身旁的人说话,AI 会误将其作为任务去应答。为了解决这一痛点,OpenAI 开始在其语音模式中引入主动反馈(Backchanneling: 在对话中发出微小声音以示倾听),而 NVIDIA 的研究模型 Personal Plex 已经实现了流畅的实时打断(Interruption)与话题恢复。这种主动参与的能力在团队协作中更为显著:在多人讨论中,AI 能够通过语境自动识别多方表达的属性(如提案、疑问、答复),在无人发言时选择合适的时机主动介入,并自动记录会议纪要和澄清设计边界。
Original English Source
Okay, let's make this concrete and familiar. A few weeks ago, my co-founder was using a Frontier company's voice mode. These are known as speech-to-speech models. He asked it a normal question, "When is the next Timberwolves game?" Fine. It answered it quickly. Then, he pretended I showed up as if to speak to me and said, "Hey Ted, come on in." He wasn't talking to the AI. But these models have no way to know that. So the AI did the only thing a prompt box can do. It took his speech as a turn and answered it. "Sure, I'm here. What's on your mind?" That's not a good answer, but it's not a dumb model. It answered the first question perfectly, but it's a protocol with exactly one slot. Your message, then it's reply. It has no concept of who's speaking, whether the words were even meant for it. And the frontier companies want to make strides as well. OpenAI released GPT real-time 2 in late May and started trying it for their voice mode as well. It backchannels now. It goes, "Mhm." and right. The little sounds we make to show we're listening actively. We're seeing the field is converging on the same conclusion we built our company on. The interface has to stop being batch and start participating. Others are working on real-time conversation as well. This is Nvidia's Personal Plex, a research model, not ours. Watch what happens when it gets interrupted. >> I've been thinking about starting a diet. >> Yeah, starting a diet can feel a bit daunting, but you could keep it simple. Focus on eating more veggies and fruits. Try to >> Before I forget, I signed up for a marathon. >> All right, congrats on signing up for the marathon. That's a big challenge. You've got a lot of time. Focus on building a solid base with regular long runs. Stay hydrated. Make sure you fuel right before and after, and don't forget to stretch and take care of your feet. >> Personal Plex stops. It yields. It picks the thread back up. That's real turn taking, listening and speaking at once in real time. >> You need to come visit me. >> Oh, okay. >> go into the city. >> Okay. >> Cuz that's the thing, like there's like the random spray paint, but then there's also like I'm not sure. People must commission them. Like these massive like mural spray paint pieces. >> Yeah, I think they do. >> And Persona Plex's back channels listens and lands where a person's would. Beyond conversational flow, there are lots of challenges and it's a complex problem. Making listening noises is not really the same as knowing who's in the room. These are not trained to tell that "Hey Ted" wasn't meant for it. But we are. We're working on improving the protocol to the models. By giving it a better understanding of human and group conversation. >> Good afternoon, everyone. >> Good afternoon, Sam. >> Hi, Jordan. >> Good afternoon. Good to see you both. >> Quick one, which requirement is this? Do we have an ID? >> This is our Q442, expense approvals. >> There, it answered a question. It only takes actions based on the utility-driven model. So, it creates goals to fulfill as it labels each of the participants' statements as a question, a proposal, an answer, and then only takes a turn when no one else is speaking or holding the floor. >> Right, we need users to approve requests faster. >> Yeah. The approval flow's too slow. >> What kind of requests, though? >> Expense approvals first. Access requests eventually. >> AI, hold that. >> Actually, let's pause. Expense approvals or a general approval workflow? >> Expense approvals first release. >> Access requests are future scope. >> That changes the data model. Good to know. >> AI, pull that up for everyone. >> Tracking determining who's the speaker referring to is critical. In this case, it was easy with a direct reference to the AI, but it will happen again without a direct reference. >> Oh, I'd forgotten that was a rule. >> So, over 5,000 needs a second approver? >> Yep. Manager plus finance. >> So, a big one can't be a single tap. >> Right. Over the limit, it routes to a second approver. >> Agreed. Under five, one tap's fine. >> Works for me. >> Okay. Agreed. Expense approvals, 5,000 threshold. >> Right there, the AI resolved the scope objective. No one wrote the prompt, no one patched the turn and hit submit. The system was in the conversation, following it, understanding, and choosing its moment.
界面技术的范式转换与消除“交互税”
从底层来看,AI 不仅仅是一项智能技术,它正日益演变成一种颠覆性的界面技术(Interface Technology)。未来的设计重心不应是让机器显得更聪明,而是去思考:我们是否仍在让用户承担不必要的认知负担?
在人机交互的演进历程中,从打孔卡、命令行、图形菜单,到触屏与提示词输入,每一次进步都伴随着沉重的交互税(Interface Tax):
- 转译税(Translation Tax):将模糊的想法转译为机器能懂的指令。
- 精确度税(Precision Tax):必须用极其精确的结构描述意图。
- 上下文税(Context Tax):手动拼凑和提供背景信息。
- 修复税(Repair Tax):不断调整输入以修正错误的输出。
AI 交互的终极目标应当是让这些负担彻底消失。理想的界面应该采用人类天然的沟通媒介——一个眼神、一段静默、一次草图或一份清单,由 AI 自动在最合适的时刻选择最匹配的物理通道与反馈方式。当机器能够顺应人类,而非强求人类顺应机器时,交互阻力将彻底消散,人机协同的真正潜力才得以释放。
Original English Source
AI, capture that for us. >> First release supports expense approvals only. Access requests are out of scope. Managers can approve or reject an expense right from a notification. And per the finance controls policy, anything over $5,000 routes to a second approver? >> Actually, make the threshold 10,000, not five. >> Want me to update the requirement to a 10,000 threshold? >> Yes. >> AI, is this room free after the meeting? >> Let me check. The room looks free until 3:00, but yes, it's yours until 3:00. >> And that's the difference between a smart machine behind the same old prompt and an interface that finally participates. Here's the mindset shift I want to leave you with. AI is not just an intelligence technology. It's increasingly becoming an interface technology. And if so, then book smart models alone are not enough. Stop picturing AI as a smarter machine hiding behind prompts, agents, loops, and all the old paradigms. We have to start seeing intelligence itself as a thing that can finally remove interface constraints and amplify human potential. For 75 years, humans adapted to the machine. It's syntax, it's forms, it's timing, it's batch. A system that can reason, listen, infer, adapt should be able to meet us partway, if not all the way. Instead, if AI is for users, then we should obsess about maximizing the interface. So, then the design question changes. It needs to become what burden are we still putting on humans only because the machine used to be too limited to carry that burden itself. Ask that question, and the whole interface space opens up. The answer isn't always chat. It isn't always voice, and not a wall of markdown. It's definitely not a decade-old set of digital constructs. The right answer is the affordance humans already use with each other. Communication, a question, a pause, a sketch, a checklist, a quiet aside, or saying nothing at all. An interface where timing and modality aren't the humans' job anymore. Where choosing the right channel at the right moment is done by the AI. And as a usability person, this is the part that excites me the most. When you take that burden off people, the the disappears and adoption follows. Computing has mostly been about to date improving how humans encode their intent for machines. The punch card, type a command, click a menu, use your thumb on an iPhone, write a prompt. Every step was progress and every step carried the old constraint forward into the next era. A translation tax, a precision tax, context tax, repair tax. AI is our chance to put those down. Not by making everything magical. Not by making everything voice. Not by replacing human judgment, but by making computers for once more fluent with us. Most talks and videos cover how to use or adopt AI. The deeper question is how AI intelligence changes the interface. Human conversation is the most human thing we do. Because if a machine can finally understand more of what we mean, then we can and should stop reshaping ourselves to be understood by it. Thank you.
📌 文中提及的人物和组织
公司/组织: JoinIn AI, OpenAI, NVIDIA
产品/模型: ChatGPT, Personal Plex