知识认知鸿沟与 AI 的体验困境
在软件工程领域,开发者与最终用户之间天然存在着一道信息差。软件工程师的技术素养普遍高于普通用户,而用户体验(UX: User Experience)正是消除这一知识鸿沟、将开发设计初衷转化为用户直观感受的桥梁。如果用户因为厌恶交互过程而拒绝使用,那么软件功能的强大将变得毫无意义。
当前,AI 界面设计面临着严峻的体验问题。各大软件商正迫不及待地将 AI 功能强推到从办公软件、社交应用到手机操作系统的各个角落。然而,这种急躁的技术堆砌往往忽略了用户的实际接受度。许多 poorly designed 的 AI 功能被直接推给用户,当用户尝试使用并得到不符合预期的垃圾输出时,沮丧和失望感便会转化为对 AI 技术的抗拒。由于第一印象被破坏,用户会逐步放弃尝试,这使得开发者在失去用户信任前只有极少数的试错机会。
Original English Source
Hi. My name is Kat. I'm a senior design and developer advocate at Progress Software, and I have spent most of my career focused on the front end. After taking a somewhat meandering path through design and into development, I eventually found my niche in the space where those two worlds overlap. My design background has given me a lot of insight into the UX side of software engineering, and the user research that I get to do as part of my work as developer advocate is one of my favorite parts of the world. I have spent a lot of my career teaching developers about design, about UX, and about how to build more user-friendly software. Something that's always been valuable, but is more crucial than ever now that AI has entered the picture. I don't need to tell you that AI is a pretty unique new technology. It is rapidly developing and has so much potential to improve systems and software in ways that we couldn't have even dreamed of just a few years ago. However, along with all of that potential is also quite a bit of risk. AI's non-deterministic nature means that the output can wildly range in quality, format, style, and accuracy, even when we're trying our best to prevent that. On top of that, it requires entirely new types of interaction patterns to use. In fact, we've seen the emergence of new vocabulary, techniques, and sometimes entire roles meant to address that need. For those of us that are living and breathing this technology every day, concepts like prompt engineers, hallucinations, retrieval augmented generation, and so much more have all become regular parts of our lives. But for the vast majority of our users, that simply isn't the case yet. In any kind of technology, there's a gap between the developer, who has by nature kind of become the subject matter expert, and the user. The technical literacy of the average software engineer is always going to be higher than the average user of said software. That's where UX comes in. It's how we bridge the gap between what we intended when we built an application and what the user actually experiences. After all, it doesn't actually matter what our software can do if our users hate using it so much that they will avoid it at all costs. And right now, AI has a kind of serious UX problem. Our users are seeing AI features get added to basically everything right now, from their work software to their social media apps, right down to their computer and phone operating systems. But, as a wise technologist once said, "Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should." We're really pushing our users to adopt these new tools and features, but we're not always making it easy to do so. It's common to see situations where a user is presented with a poorly designed AI feature that they really only have the barest idea how to use. Then, when they try it and it doesn't do what they wanted or expected it to do, they feel annoyed and off-put and kind of negative about AI generally. The more times they try an AI feature and get sub-par results, the less likely they are to engage with it again in the future. Which means that, as the creators of these features and software, we really only have so many chances to get this right before our users just disengage and stop trying at all.在确立了用户体验对缩短技术鸿沟的必要性后,我们需要从历史中寻找灵感,了解如何将这种新技术平滑地引入用户的日常工作。
借鉴 Macintosh:心智模型的移用与对话模式重构
当 Macintosh 首次引入图形用户界面(GUI)时,对当时的用户来说,它和如今的 AI 一样陌生。Macintosh 成功的关键在于其卓越的 UX 策略:借用现实世界中的词汇和工作流,使用直观的图标和视觉元素,引导用户逐步熟悉个人电脑。随着时间推移,用户数字素养提升,界面才逐渐舍弃了如“兔子与乌龟”等具象的隐喻,引入了更专业的术语。当前 AI 交互的发展大约仅相当于 Macintosh 系统的 System 3 阶段——我们虽然不需再依赖过度简化的隐喻,但也不能预设用户都是 AI 专家。
幸运的是,用户在传统软件中积累的心智模型(Mental Model: 用户对系统如何运作的内部假设)可以直接复用。例如,在 AI 聊天界面中,我们可以保留双向对话气泡、消息历史记录和发送按钮等熟悉的 IM 设计模式,作为用户的切入点。然而,人机对话与人际对话存在本质区别。普通的聊天界面无法承载诸如添加参考源、随时中止 LLM 回复或调用智能体工具等特有需求。以 Microsoft Teams 与微软 Copilot 研究智能体为例,Teams 聊天默认用户具备高成熟度,隐藏了大量图标释义;而 Copilot 则通过丰富的提示示例和清晰的语音选项,极大地降低了准入门槛。
Original English Source
In my personal opinion, this knowledge gap between developer and user when it comes to AI is one of the widest that I've seen with any technology. That makes it extremely hard for us to build AI-powered features that our users are actually going to get the full benefit from. When Macintosh introduced their graphical user interface, it was, like AI is now, something that was very different and unfamiliar to their users. But a large part of their success and reputation was built on their ability to create a user experience that met users where they were. They borrowed vocabulary and user flows from real life. They included lots of icons and visual representations, and they really made sure to walk users through how to get the most out of their new personal computers. And over time, as the average user became more comfortable and more literate in the space, the user experience of the Macintosh evolved with them. We start to see more technical terms and fewer abstractions. Look at the difference here between the control panels in Macintosh system 1 and system 6. If you had presented a system 1 user with the words rate of insertion point blinking or RAM cache, it would have been pretty much meaningless to them. Similarly, we see that those low and high volume icons are gone in system 6, along with the little turtle and rabbit speed icons, right? That kind of more literal depiction wasn't needed anymore. Right now, we are probably somewhere around system 3 in this metaphor of introducing AI to users. We may not need the turtle and rabbit icons anymore, but we also can't yet assume that they will sit down with our AI software as experts. The good news here is that, unlike the Macintosh UI, we're not totally starting from scratch here. Our users have existing mental models about software usage that's going to carry over to what we're building. They're just not always going to map over exactly one to one. But that's where we come in. As developers, it's our role to leverage these patterns that our users are already familiar with, start to combine them in new ways, and then introduce them to our users gradually in a way that's not going to feel overwhelming. A good user experience will not only provide a user with the tools, but also guide them how to use them. For instance, let's look at an AI chat interface. From a purely UI perspective, chatting with an LLM is almost, but not quite, like chatting with another user. So, we'll be able to borrow some familiar patterns here, like having the user's chats show up on one side and the AI chats on the other, messages that appear above an input field with a send button, a scrollable message history, and more. That gives our users a really good jumping off point and a lot of visual clues about how to start interacting with an AI chat. However, the experience isn't quite close enough that we can just lift an existing interface from direct messages and repurpose it wholesale. A chat interface that was originally built for two people is not going to have the accommodations that we need for things like adding reference sources, pausing or stopping an in-progress reply from the LLM, leveraging agentic tools and integrations, or more. So, those are going to be the places where we have to adapt existing UX and UI patterns into something new. We have to be the bridge and help our users transition into familiarity with these new AI experiences. We can see the difference here, right? Compared this Copilot research agent new chat with the new Teams chat. Both of them have familiar aspects, right? We see text boxes, we see the plus to attach files, and so on. But, the Teams chat assumes a much higher level of familiarity. It's not spelling out any of the iconography it uses, and we don't see some of the things intended to lower the barrier of entry, such as the examples, large text buttons, or voice interaction options that show up on the agent chat.在明确了不能直接搬用旧界面的局限后,我们必须探讨另一个关键争议:为何不直接用 AI 来自动生成用户界面?
为什么 AI 无法代替人来设计 AI 界面?
面对全新的交互挑战,一个常见的疑问是:“为什么不直接让 AI 自动生成 UI 界面呢?”然而,AI 只能对已有的模式进行重新混合(Remix),而关于 AI 专属界面的设计范式目前正处于快速演变中,互联网上并不存在现成的“最优解库”供其参考。过度依赖 AI 自动生成界面只会导致最终结果高度趋同和趋于平庸。
人机协同的设计过程必须由人类开发者与设计师来主导。基于大量的用户研究与测试,我们可以将用户在使用 AI 时面临的挑战归纳为五个核心范畴:信任(Trust)、清晰(Clarity)、控制(Control)、透明(Transparency)以及实际价值(Meaningful Benefit)。
Original English Source
Now, before we really dive into this, I do think we have to take a moment to address kind of the elephant in the room, right? Why not just have AI create these solutions for us? Why do we have to be the ones who design and build new AI patterns when we have technology that can generate interfaces for us now? The problem with that is that AI can only really remix things that already exist. It's fantastic for looking at those existing common patterns and replicating them, but the patterns for AI interfaces don't fully exist yet. Or at the very least, they are still in the process of being rapidly developed, and they are changing all the time as that technology advances. There's just not a long history of standardized and familiar AI-related interfaces that we can reference to solve this problem. So, at least for now, we're kind of on our own here, and we can't yet defer it to AI. Maybe that will change down the road, but for the time being, this is still a very human problem. Additionally, as many folks have begun to notice, the design that an AI creates is going to be a lot like every other design out there. AI can reference and remix, but you're just not going to get brand new concepts from it. An AI-generated UI tends to look pretty darn average. And while that can be a really great starting point, it's not usually a great ending point. If you just need a quick landing page, it's probably going to do the trick. But since we're dealing with whole new interaction modes that need more attention paid to the UX rather than less, it's just not going to be enough to get us where we need to go right now at the quality level that our users need and deserve. So, if we're going to build new patterns and new interfaces to help our users get the most value out of this technology, where do we start? Well, like most things in design, it's best to start with the user. Specifically, the user's problems. That means that when we're thinking about building our AI features, we need to be thinking about what challenges our users have with AI and the user flows that will mitigate those issues as much as possible. From the user research I've been doing and the users I've had the chance to speak to, I've started to group the main challenges here into five categories. Trust, clarity, control, transparency, and meaningful benefit. In this talk, we're going to discuss each of these in depth, understanding the problems our users have when they're using AI, and looking at some examples of new patterns and techniques we can employ to address them. By making sure that we are addressing each of these, we can start to create AI experiences that support our users through the introduction of this new technology and ensure that they are in a place to use it to its fullest potential. So, let's get going. >> [snorts]在确立了五大核心维度后,我们首先需要解决的是阻碍用户接受 AI 功能的最大拦路虎——信任危机。
核心支柱一:通过可验证性与“人机协同”建立信任
对普通用户而言,AI 是一个不可预测的“黑盒”。由于大语言模型无法 100% 避免幻觉(Hallucination: 模型生成看似合理但与事实不符的信息的现象),宣称“我们的 AI 比别人更安全可靠”根本无法说服用户。要让用户信任系统,我们必须践行“信任但验证(Trust, but verify)”的策略,为用户提供自我验证的工具。
- 信源引用: 在 AI 生成的响应中明确标注引用来源。对于简短的背景参考,可以使用工具提示(Tooltips: 悬停显示的轻量文本框);对于互联网公开内容,提供行内链接;对于复杂的学术或行业研究,则推荐在右侧拉出侧边栏参考窗格(Side Panel Reference Windows),将 AI 重新定位为“信息图书管理员”而非权威的“全知专家”。这使用户在将生成内容用于工作报告、邮件等场景时,免于承担因事实性错误而导致的个人信誉受损风险。
- 人机协同智能体流: 在自主思考和执行任务的智能体工作流(Agentic Workflows)中,当涉及高风险操作时,必须展示一个明确的行动计划(Action Plan)并等待用户批准方可执行。虽然对于重复工作用户可以设置一键免签,但在首次运行或高风险环节中,人机协同(Human-in-the-Loop)机制依然是不可或缺的信任基石。
Original English Source
Trust. The number one, by far, biggest hurdle that we need to overcome in order for our users to engage with the AI features we build is trust. Right now, to the average user, AI is a black box. Many simply do not understand from a technical perspective how it works. And when you don't understand how something works, it makes it very, very hard to trust the output. On top of this, just about everyone who has interacted with an LLM before has, at some point, seen them hallucinate. Even as the models keep getting better, we're still not at the point where we can claim any tool will be 100% guaranteed hallucination-free. There's always a chance, even if it's a small one, that the AI is going to provide an incorrect response. So, how many times can a user see output that's wrong and still trust it? It can be easy to think that the answer to this problem is to try and position the feature that we're launching as some kind of exception to this rule. We say things like, "Other AI tools may not be trustworthy, but ours is different. Ours is higher quality, it's safer, it's more reliable." So on and so forth. But not only is this questionably true, right? I mean, after all, and not many of us are really training our own models here. So a lot of this is going to be simply outside our control. But it's also a very difficult thing to try and sell your users on. In situations where we cannot promise the truth, we have to go above and beyond to earn it. Our honesty about current capabilities and limitations of our AI tools will go a lot further with our users than the denial of any potential problems. So rather than trying to convince users that our solution is inherently trustworthy, we can instead build in patterns that empower them to see when incorrect information is returned and give the tools to correct or mitigate it. You know the saying, "Trust, but verify"? That's kind of the goal here. By citing the references in an AI-generated response and linking users back directly to the source material, we give users the information they need to validate AI output themselves. The more a user is able to click through and see where an answer came from, the more they'll be able to trust the content, even if they don't choose to check every source every time. In addition to allowing our users to vet the answers, citations also have a second important purpose. They allow users to repurpose the material in their own work while maintaining a trail of accuracy. This is really important for their own credibility and reputation. After all, how often would you share content from an unverifiable source if you knew that any errors would ultimately be attached back to your own name? Generating output for the user might be the last step in our process, but it's really just the beginning for them. Users want to take that content and turn it into a report, an email, a campaign, a presentation. And if the output can't be cited and trusted, then it's just not going to be used. There are a handful of different ways we can implement this, and which one you pick is going to depend on what you're building and the context in which it's being used. But a few popular approaches include things like tooltips, inline links, and side panel reference windows. Tooltips allow you to provide a small snippet of a relevant quote, which can be really great for increasing credibility. Links, of course, are ideal if the content is coming from an external source, like another webpage, as opposed to an internal source, like a document in a shared drive. And if the feature you're building is meant to support more research-oriented work, then you might consider adding a side panel, where the user can explore the source material in more detail, which kind of positions the AI assistant as more of a librarian than inherently a subject matter expert. The other place where trust factors heavily into AI work is agentic workflows, where the AI is thinking and executing work on its own. This, understandably, can be kind of concerning for users, depending on the stakes of the project and how easy or hard it would be to roll back any incorrect actions, which is when keeping that human in the loop really becomes crucial. One of the best ways to help users trust these systems enough to use them is to show an action plan to the user and allow them to approve it before the agent begins work. This has become a pretty common flow in a lot of foundation models, right? If you're using agentic features in Claude or ChatGPT, you will have seen it create a list of the steps it will take to accomplish a given task and then ask for your confirmation before beginning. Of course, there are also settings you can add to toggle this off or always allow it, all of which is important if you're going to allow users to repeat a flow over and over and don't want them to have to baby sit it. Spoiler alert, we will talk a little bit more about permissions in the transparency section later. But if you're creating some kind of an agentic tool and no plan is ever shown to the user and an action happens without them understanding why or how, it can be very, very hard for them to trust both the result of that action and the agentic tool itself.在通过信源和流程确认建立信任之后,我们需要解决如何明确标识 AI 的边界,让其褪去神秘色彩。
核心支柱二:去魔化与系统状态的“清晰”传达
用户对 AI 生成的内容普遍存在偏见。一部分人极其反感并将其斥为“毫无价值的垃圾(Slop)”,另一部分人则感到被欺骗。要解决这种不安感,我们必须确保清晰度(Clarity)。首先,绝不能将 AI 包装为一种神秘的“魔法”——目前软件中泛滥的“星星”或“闪烁(Sparkle)”图标就是这种神秘化的反面典型。AI 本质上只是另一种技术,用户应当拥有知情权。
我们必须在界面中明确标识哪些内容是 AI 生成的,避免用户混淆人工作品与机器输出。除了基础的AI 生成标记(AI-Generated Tags),我们还可以引入系统级注解(Systems Lens)——如列出当前输出的局限性提示,甚至提供由模型自评估的“可信度评分(Confidence Score)”。这种反向兜底设计虽然坦承了软件的瑕疵,但却向用户提供了透明的运行状况,有利于建立更长久的用户粘性。
同时,必须做好系统状态的传达。当 AI 在后台运行庞大的计算任务时,界面不应当只显示一个静态的无意义加载器。理想的设计是采用渐进式的文本更新或交互式步骤,清晰告诉用户 AI 正处于“哪一步”:比如“正在搜索内网...”,“正在汇总报表...”。当生成的内容插入页面时,界面应通过柔和的动画平滑推移已有布局,而非突兀地导致屏幕内容跳动(Layout Shift),这能帮助用户保持注意力的聚焦。
Original English Source
There is also, undeniably and I would maybe even argue correctly, quite a bit of skepticism around whether or not content is AI-generated right now. You've probably seen exchanges online where someone shares a photo or video and another person is immediately commenting to tell them that's not actually real. When we spend so much of our time and energy second-guessing and investigating content that's shared with us, it doesn't exactly foster an environment of trust. For now, AI-generated content is highly polarizing with users. While some have heavily leaned into using it in their daily lives, others will react to it negatively and dismiss it as slop, right? The ability for AI to generate text, images, and video that is nearly indistinguishable from human-created work is very new. And for many users, that still feels uncanny and unsettling. I don't say this to start any kind of debate, right? I mean, seems harsh to say it, but our own personal opinions on AI-generated content don't really matter too much in this context. What does matter is the fact that our users feel this way, and they deserve to know what kind of software they are using. The answer here is clarity. We have to be honest with our users when they are interacting with AI-generated content. And the absolute worst thing we can do when we want to introduce clarity to a system is to present the same AI to our users as magic, right? We don't even have to look further than the prevalence of the sparkle icon to designate AI as an example of that. However, while that might feel kind of mysterious and cool, it's really not accurate, right? AI is just another technology, and our users deserve clarity into when and where they're interacting with it. Visual cues like clear labels or icons that specify content as AI-generated are a great first step, and providing systemic explanations or annotations that help user contextualize results will go even further. For example, if you can show a user confidence scores or direct references to resources, that's going to help them feel like they're looking at a real technology and not a black box. Clarity also plays a big role in system status. When an AI tool is thinking or generating content, we want to make sure we are showing exactly what's happening. And rather than just using a generic loading spinner, we can use visual cues to show progress, such as generating text in real time or showing a list of steps in a multi-step process as they are completed. We also want to pay attention to how that generated content is inserted onto the page. For example, if you have a page that is dynamically shifting up or down to accommodate new content, that can be a really jarring experience for users who are trying to read it. In standard UX, we use animation to bridge that gap. We can use transitions to slide content up or down the page to see what's changed, right? Like when new messages appear in an ongoing chat and the message history scrolls down to display it. Other times though, it could just mean drawing the user's eye back to a space they might not have been actively watching. Even if that doesn't include the scrolling, just a subtle animation or a highlight can help draw focus without interrupting their flow.当系统状态能够被清晰感知后,交互的核心将转移到如何赋予用户对生成逻辑的掌控权。
核心支柱三:紧急制动、安全探索与精细化控制
非确定性的输出意味着用户在协同过程中必须拥有绝对的控制权。任何在生成中不允许中断、不能回退的设计都是不可接受的。
- 紧急制动按钮: 界面上应当提供一个极其醒目的“一键制动(Emergency Brake)”按钮。在智能体自主写入文件、运行脚本、操作浏览器等高风险流中,用户必须能够在察觉异常的第一时间强行中止进程。
- 安全探索: 引入精细的版本历史机制。传统的撤销(Undo/Redo)功能通常支持撤回最近几步。对于涉及长文或复杂文档编辑的 AI 助手,我们需要引入保存状态(Save States)或检查点(Checkpoints)机制。
- 定向调整: 允许用户对 AI 输出的部分段落或选区进行微调(Targeted Adjustments),而非强迫用户在对部分结果不满时点击“重新生成”,导致已满意的部分随之丢失,并造成昂贵的令牌(Tokens: 大语言模型处理文本的计费单位)浪费。
Original English Source
When users are working with our AI tools, they need to know that they are still ultimately the ones in the driver's seat. After all, there's very little value in being able to assign the AI tasks if you can't then step in as you like to make adjustments, corrections, and fixes. In that vein, it is also important for the user to be able to stop or override an AI action at any point, whether that's during content generation, partway through an agentic workflow, when it's running code or scripts, or something else entirely. If the user is unable to abort the process, then they're not actually the one in control, and frankly, that's not really acceptable. One of the most reassuring things we can offer nervous users is a clear and prominent emergency brake that they can slam to bring everything to a halt for any reason. That means it can't be tucked away in a menu somewhere or involving some specific command they have to remember, which they won't in a high-stress moment. We want to give them the equivalent of a big red button they can slam to make everything stop whenever they need. Of course, once a user has stopped a process, we can probably guess what they want to do next, right? They want to undo the incorrect things that changed. Even way before AI, Nielsen's heuristics list safe exploration as one of the core principles required for a good user experience. That means that users should be able to navigate back and forth, click and unclick things, and generally just kind of mess around in a piece of software without getting stuck or unintentionally making big permanent changes. Sometimes a user will try something and then change their mind, decide they don't like it, or it didn't do what they thought it would. In those cases, we want to make it as easy as possible for them to roll things back to an earlier state. This kind of safety net was always valuable, but now that we're dealing with non-deterministic output, having some kind of a version history is going to be pretty much a non-negotiable. Of course, the level of version history that's required will differ based on the task. The Git style of version control that we are used to as developers is probably more than you will need in a user-facing application. If this is a truly simple conversational interface, you may not need it at all. The user could just rephrase their question and try again if they don't get the answer they wanted. However, if your tool is doing something more complex, like creating and editing a document, simple undo and redo actions that allow users to move back and forth within the last 10 or so steps would be really helpful. And if you're wanting them to do truly advanced work with your tools, then it's worth building in the mechanisms for checkpoints or save states. Consider what kind of tasks your user is going to be working on with your AI features, and think about how significantly each step forward is likely to change the existing content. Ideally, this will also allow users the specificity to make targeted targeted adjustments. You don't necessarily want them to have to wipe out and replace everything that happened in a given step if they weren't happy with it. It's common for AI-generated results to have a mix of quality in the output. Users may want to keep some aspects while reverting others. And the more control we can give them over that creative process, the more enjoyable it will be for them to collaborate with our AI tools. This is especially true now as the cost of tokens is starting to tick up. Having granular control over what gets reworked allows for more specific and production productive iteration without them just having to say, "Try again." and start over from scratch each time.在处理了细粒度交互控制后,我们必须向上探究另一个更宏观的问题——用户数据管理和操作代理中的权限透明度。
核心支柱四:数据透明与权限的灰色地带
当 AI 系统获得读取本地文件、发送电子邮件或执行脚本的权限时,传统的“是否同意”二元复选框已无法满足安全诉求。透明度(Transparency)要求我们必须能够清晰展示 AI “正在用什么数据做什么事”。
- 权限分级机制: 开发者必须识别权限授权的“灰色地带”。用户可能只允许 AI 引用某些数据库,但决不允许其执行修改或删除表的操作;只允许它读取邮件,而不能以用户身份发送邮件;只允许读写某个特定的临时文件夹,而非系统根目录。
- 记忆主权与成本管理: 如果 AI 系统具有记忆机制,必须提供明确的“被遗忘权”,允许用户在管理后台中查看 AI 记住了哪些偏好信息,并允许其随时一键清空或局部遗忘。此外,在发起耗时较长或需要调用昂贵外部服务的任务前,界面必须提供预估时长和预估 Token 成本。
- 独立运行状态指示: 当智能体脱离用户视角开始独立工作(例如接管用户浏览器、在后台多步爬取网页)时,界面必须以显眼的形式(如全局通知横幅、醒目的彩色外框或锁定罩层)传达这一状态。不仅为了告知用户系统正在自主运行,更能防止用户在不知情的情况下干扰或阻断正在执行的复杂自动化流程。
Original English Source
AI tools are strongest when they seamlessly integrate into the rest of a user's workflow. Referencing internal documents, sending emails, making calendar appointments, drafting and sharing content, so on and so forth. However, most users won't integrate what they can't see and don't understand. If the AI applications we build are going to become part of their daily lives, then users need transparency into exactly what they do, what they have access to, and what it will cost them in both time and money. The more transparent we can make these features, the less hesitation our users will have in adopting them. The concept of asking for user permissions is certainly not AI specific, but the stakes can feel a lot higher when you're asking users for permission to allow AI to act on their behalf. That means we need to make it easy for users to see which applications we've allowed our tool access to and in what ways, keeping in mind that permissions are often more than just a binary yes or no. So, can your agent just reference data from this database? Or is it also allowed to delete tables? Can only read your user's emails, or can it send from their addresses well? Can it run scripts, search their local files? Each one of these actions requires direct user sign-off. Permissions are also not a one-and-done situation. A user might feel comfortable allowing an action to happen once under their direct supervision, but they might not want to allow it permanently. They could give access to a specific folder, but not every folder. They might approve an action, but then want to be notified every time it's taken. Consider which of these gray zones might exist in your permission structure and try to accommodate as many different options as you can. Another common sticking point with AI is data collection. Often, it can be helpful to save information from past interactions, but storing this kind of data requires real attention to user permissions and data management. If you want to create a system that can remember things, then you also need to make sure that it can forget and that the user can not only see, but has the final say in what exactly gets remembered. An ideal permissions flow will not only ask for a user's approval in the moment, but also create a space where they can see the history of what they've given access to and when, as well as allowing them to revoke that access or permission at any time. Now, this one's pretty simple, so we won't spend too much time on it, but another crucial aspect of transparency is making sure users know exactly what they are committing to when they approve an action request. In addition to approving those actual steps and plan, as we already discussed in the trust section, they also need to know how much time something is going to take and what it's going to cost them in money or tokens or credits, however you are counting this. Even if we can't specify an exact amount, if we can provide a rough estimate, that's generally going to give users enough information to work with and to potentially revise their request if it's going to exceed what they're comfortable with. Similarly to the marking that denotes which content was AI-generated, it's also a really good example uh of an chance to have a visual signal for when an AI tool is acting independently. If, for example, you're going to allow an agent to take control of the user's browser, then there needs to be some kind of banner or sidebar or outline, an indication that the user is no longer driving that interaction. Not only is this a good thing to do just for transparency, so the user understands the current state of the system, but it's also helpful to make sure they don't unintentionally interrupt or confuse an ongoing process. It's an especially important consideration for processes that you know are going to take an extended period of time where the user might have stepped away and then come back without necessarily keeping track of everything going on while they were gone.在解决了安全和透明度后,我们最终需要回归到 AI 的核心目的上:如何切实降低用户使用门槛,交付有价值的业务产出。
核心支柱五:告别“空输入框”与重塑工作流集成
如果 AI 无法帮助用户解决实际的生产力痛点,那么再炫酷的技术也只是暂时的“新鲜玩意”。要实现实际价值(Meaningful Benefit),开发者必须在两端进行体验优化:
- 消除“空白文本框”: 放置一个“Ask AI”的空白输入框,实际上是将“如何编写提示词(Prompt Engineering: 包含指定上下文、输出格式规范及拆分复杂任务的技巧)”的认知负荷强加给了初学者。为了降低门槛,我们必须在空白界面中提供场景模板(Templates)、示例提示词(Suggested Prompts)和引导式工作流(Guided Workflows),演示系统的能力边界。
- 一键集成下游任务: AI 生成内容只是工作的起点,而不是终点。当系统生成数据或图表后,必须提供下一步操作按钮(Next-step Action Buttons),引导用户快速将其应用到实际业务中。更好的方式是建立工具直连集成(Direct Integrations):例如支持将生成的文本草稿一键推送到 Microsoft Word 或 Google Docs,或者将漏洞诊断结果一键同步为 Jira 或 GitHub Issues 中的故障工单。
这正是当前 AI 应用的分水岭。随着底层基础模型能力迅速同质化,决定一款 AI 软件生死存亡的要素已不再是模型跑得多快,而是围绕模型构建的交互体验质量。将用户置于设计的核心,为他们保留控制力、知情权和尊严,才是打造用户不反感的 AI 应用的唯一解法。
Original English Source
Finally, all of the interesting technology and cool functionality in the world is meaningless if what we build isn't solving the problems that our users need solved. To do that, we need to create experiences that make it as simple as possible for our users to get the output they need from our AI features and start to leverage it in their work or daily lives. One of the biggest mistakes we can make when developing AI features is assuming that our users know how to use them. As mentioned at the beginning of this talk, right? We know how to do things like provide context or specify formatting requirements or break large tasks into smaller ones and then iterate, but our users often don't. So, when we place a blank text box in front of a user and just tell them to ask AI, we're actually kind of asking them to do a lot of work in figuring out how to really use it. They need to understand what kinds of problems the tool can solve, how specific they need to be, what information the AI tool needs to succeed, if they need to add sources or references, how to recognize when a response needs refinement. Rather than expecting this level of AI literacy from our users on day one, we can help them by providing examples, templates, suggested prompts, and guided workflows that demonstrate what success looks like. Now, what will your users be doing with the content that your AI feature generates? If they're going to run a search or analyze a spreadsheet, what happens next? What do they do with the results? If they create an image, who are they showing it to? How are they sharing it? Is it getting printed or posted online or sent to a friend? By introducing next step action buttons, we can make it as easy as possible for them to leverage the output of our tools in their own work. If they're not able to action on the content that was created, it's never going to be more than a novelty to them. It's not going to be of realistic use. But we can help guide them to suggested next actions that will help them make the most of what they created. If we want to take this one step further, we can start to build in direct integrations with their other most used tools. Like, maybe they want to generate a new document with the content in their word processing software from a draft. Maybe they want to push new code to a linked repo. Maybe they want to open a new ticket in a tracking system with their findings from a test. The simpler we make it for them to get the information from our tool into the rest of their workflow, the more useful it's going to be to them. Now, AI can generate content, write code, analyze data, automate tasks, empower our users to do all kinds of amazing work. But only if they're willing to try it, and only if we make it easy enough for them to do. When everything our AI tools are doing is hidden behind a curtain, it makes users feel like things are happening without their input, which is hard, especially when many of them are already kind of feeling some level of AI skepticism or hesitation. If we unintentionally create AI experiences that take away our users' power, understanding, and autonomy, then they'll never be interested in using what we build, because they'll never feel truly comfortable engaging with it. By focusing on patterns that reinforce those pillars of trust, clarity, control, transparency, and meaningful benefit, we can set guardrails around our AI features that will help make our users feel safe and confident. Cuz the technology is here, right? The models are already really good, and they just keep getting better and better, faster and more efficient. That means that the differentiator for the AI-powered software we build isn't performance anymore. It's the quality of the experiences that we can build around them. The line between design and development is blurring a little bit more every day. How users interact with the system, how much information they see, how much control they have, and how we earn their trust are now questions that developers have to consider when building AI-powered software, whether your title includes designer or not. The technology may be incredible, but for it to be truly successful, users still need to be at the center of everything we build. Thank you guys very much for listening to this talk. I really hope this is helpful to you as you start integrating AI features into your own applications and software. You can find these slides, as well as a full transcript of this talk, at the following links. And, of course, feel free to reach out to me online if you have any questions at all. I'm always happy to chat.📌 文中提及的人物和组织
公司/组织: Progress Software
产品/模型: ChatGPT, Claude, Microsoft Copilot