聊天与引用并非垂直AI终点
许多人在医疗、法律、税务等垂直行业构建 AI 产品,并向客户承诺“省时省钱,AI 代理能在你睡觉时帮你工作”。然而,当前最主流的交互界面——用于输入的聊天(Chat: 用于输入的柔性交互界面)和用于输出的引用(Citations: 用于输出和可验证性的参考依据)——并不能真正兑现这一承诺。
作为 Filed(一家专门为美国税务专业人员构建 AI 产品的公司,已融资超过 1700 万美元)的联合创始人兼 CTO,Atul Ramachandran 指出,虽然聊天界面开发迅速并打破了传统 UI 限制,引用机制也减少了幻觉并将答案锚定在事实中,但它们存在致命缺陷:
- 同步的(Synchronous: 具有实时等待特征的交互介质)聊天形式要求用户必须在平台前输入并等待响应,无法抽身去做其他工作。
- 引用则将验证成本(Verification Burden: 确保输出正确性所需的审核工作)重新转嫁给了客户,客户不得不逐条核对 AI 的工作。在高容错率门槛的垂直领域中,这种逐条核查反而增加了额外工作,导致用户抱怨 AI 未能真正解放他们。
Original English
Chat and citations won't save your vertical AI. I'm sure many of us are building products in what in a vertical industry. For it be healthcare, legal, or taxes. And usually doing so, we sell one promise to the customers. Either we're going to save you money or save you cost. And usually the pitch goes, AI agents are here. Uh they will do the work for you while you sleep. Now, primary interface that we use for this are chat and citations. To interact with any AI agent, right? Chats are usually used for inputs. Uh they allow you to be flexible. You can talk to the agents as you please. While citations are used usually used for outputs. They allow you to see the results, verify the results, and so on. Now, I'm here to tell you that citations and chat alone will not keep allow you to keep the promise that you made to the customers. Promise of, you know, saving time and money for them. So, who am I? I'm Atul. I'm the CTO and co-founder at Filed. Uh I've been fortunate enough to build products for more than a decade now. Um at Filed, uh we build products for tax professionals in the US. We have raised more than $17 million to do so. And about 2 years now doing this, we have seen massive growth in the space. Uh just to give you an idea, last just last month, we closed more revenue than what we have done in the one year alone before that. So, the growth has been tremendous and we're very excited for this. Doing this entire journey for the last 2 years, we've learned quite a bit, you know. Uh And I think most of the learnings of building these AI agents for the taxes industry is essentially transferable to any other AI product. Now, chat is great. Like it allows us to build fast. It allows customers to interact with products in ways that was previously possible. It's a communication platform, right? Uh you can interact with agents to do tasks, flexible tasks that were not constrained to the UIs that you have built. Uh citations are also amazing. They ground answers in truth. Other than allowing the users to verify the outputs, they also serve another purpose. They allow the agents to essentially be more accurate because now they are forced to present the references to their answers. Agents perform much better with So, they reduce hallucinations essentially. But this is There's a key problem here, right? Chat is synchronous. If you think about it, even when you're coding, you're Let's say you're using product code as as as long as you're typing the request. Once the request is typed, you're waiting for the agent to respond. This synchronous medium does not allow the customers to leave the platform and go and do their work. Similarly, citations also puts the verification burden back into the customer. So, think about going and reviewing the agent's work now one by one to ensure everything is correct. Especially in vertical industry space like health care, legal, and taxes, this becomes very crucial because now this this adds an extra work, and customers usually complain that no the promise of, you know, agents doing the work for me while I sleep is not really kept here.
从自助服务演进到代理委托
为了破局,我们需要回顾产品交互的三种抽象层级演进(以银行为例):
- 物理分行:客户到网点将任务委托给银行职员,此时价值创造的瓶颈在于银行的员工数量。
- 数字化转型:移动端和在线门户让用户可以自助进行交易,瓶颈转移为用户的数量。
- 代理委托(Agentic Delegation: 客户将长时间运行的任务彻底授权给 AI 代理执掌的行为):用户不再亲自使用产品,而是扮演监督者(Supervisors: 扮演监控、协调与最终把关角色而非直接操作的人员)的角色,将工作委托给 AI 代理。AI 在用户睡觉时处理任务,消除了“用户必须在线”的瓶颈,从而创造出远超以往的价值。
在此模式下,我们应该将产品设计为一条传送带(Conveyor Belt: 一种以流式处理和自动化生产为特征的系统性产品设计架构),传送带及周边基础设施是产品本身,AI 代理是传送带上的工人,而用户则是监控并随时干预的管理者。
Original English
Now, before we dive in, I want to take you through a brief history of how products have evolved. Specifically, I think there are three levels of abstractions, I would call it. So, let's imagine a bank to make it an make it easier. Let's imagine a bank. The banks used to have a physical presence before. You would go to a bank branch, you would take the money out by talking to a employee of the bank, right? And the point here was the the person who's doing the task is the employee of a particular company. So, you as a user would go in, you would delegate the task to an employee, they would do the task for you, and then, you know, come back with the result. Be it withdrawing money or seeing how much balance you have and so on. Now, this essentially meant the bottleneck was the number of users number of employees that a company had. That's the bottleneck for creating value, right? Now, as time progressed, digital transformation era came. Basically, banks became online. You could Now, the users can essentially open up the mobile app or or the online portal and can make the transaction themselves. Can actually go and see the balance himself. This was great. Because now the bottleneck moved from the number of employees a company have to number of users that the company have. The more users meant more value you can generate. Now, I think we have reached a point of another transformation layer called agentic delegation. So, no longer the users are coming in your product to use the product. I think they're coming to delegate more and more work to AI agents to do. So, the point here is um if a user comes to your product and starts delegating work, long-running work, what would happen here is there's the bottleneck of number of users also goes away. So, it's no longer than amount of value that you generate is the amount of the number of times the user have visited your platform because agents can do the work while the users have gone to sleep, you know, have delegated the task and went off. So, the bottleneck has shifted, meaning you can generate more value than ever before for your customers. Okay, so to make it easier, I You can think of a product You have agentic product as a conveyor belt, you know, and the users as the supervisors of the conveyor belt. So, think of it this way. Previously, in a conveyor belt, usually there's a number of tasks that are happening. There are workers who are doing the work for you, and there's a supervisor who is delegating the task, you know, and monitoring how the things are going. Similarly, now AI agents are your workers in that conveyor belt, while, you know, the conveyor belt and the entire infrastructure around it is your product. So, users can come in, uh which who is the supervisor, can come into the product, can start delegating tasks to the agents. And this would mean it give it would give an idea of what are the tools that you need to build in your product to make this happen. So that the agents can do the work and the supervisor or the user is, you know, pretty confident that the work is happening.
构建传送带模式的两大基石
要构建这套传送带式的 AI 产品,前两个核心支柱是:
- 任务委托:产品经理或工程师需要识别出那些用户原本需要花费数小时处理的、可重复的痛点任务,并将其交由后台长程代理(Long-running Background Agents: 能够在后台独立运行、处理耗时几小时任务的 AI 系统)处理。在 Filed 的税务工作流中,团队已明确识别出三个此类任务。
- 技能传授:垂直行业中的不同公司和开发者拥有各自的偏好和规范,仅解决 80% 到 90% 的通用问题并不够,剩下的 10% 到 20% 属于个性化长尾需求。产品必须允许 AI 代理自动捕获用户的行为习惯并自我迭代。例如,无需为用户提供单独的规则编写界面,而是通过自动监测用户的实际产品使用行为,让 AI 代理在使用中不断学习,完成定制化技能(Skills: 承载特定公司或群体独特最佳实践与行业规范的微调规则或流程)的沉淀。
Original English
All right. So, how do we build this conveyor belt? So, think of this this way. Like, when you're trying to build a product feature, think if a user wants to delegate a task instead of doing it themselves in your platform, what would that interface look like? Essentially, design for delegation, not participation. Now, to keep it more concrete, I think there are four key pieces when building a agentic product. And there are four key features or components that you need to have so that this conveyor belt that we're imagining can work. So, first thing is the easiest one. You need to find the task to delegate. The the tasks that are coming in the conveyor belt are delegatable, right? And that can generate value for users. Second is, you need ability for users to teach the agents of, you know, uh the supervisor should be able to come in teach how the work needs to be done. And lastly, uh there's no point of a conveyor belt where you cannot monitor the work yourself, right? And and of course, intervene when something goes wrong. Now, let's take delegation. Um so, delegation means, you know, coming and handing over the task to another person, right? So, think of your users coming to a platform and delegating a task or handing off task to an AI agent. So, when building for this particular piece, you as a developer or a product engineer have to find tasks that your users do that take more than a couple of hours. So, in case of taxes, um there are like three different tasks that we have identified in the tax workflow, which take more than a an hour for our users to perform. Similarly, in your industry, you can figure out, you know, which are those tasks. It's important that these tasks are repeatable or sort of repeatable with definitely applicable to each of their, you know, use cases. And these tasks are what you use to build what you call a long-running background agents. This is where the core value is, you know, because you're taking off hours of work from your users' hands. Awesome. Um so, next is, you know, uh in any industry, in any professional industry, uh like in case of coding, for example, everyone every single user have their own preferences and have have their own way of doing a task. An example is, you know, coding, for example, we every single company or every single group of developers have their own ways of dealing with best practices and their own conventions and practices that they follow. So, if you create a background agent, for example, and it does this end-to-end task of creating a product in case of coding, for example, it will produce output, sure, right? It produce It will get you most of like 80 to 90% there, but think of it this way, like that will not yet solve the problems of of the user, right? As a user, you would want it to do it the way that you do the work. So, this is where skills come into play. Skills already exist in the today's world of agent development. So, we need to ensure that your product also has skills in place to capture where you can teach your agents how your users do the work. This is the last 20% of the of the work that you need to take care of. This is where the real value is, the quirks of the work, you know, that you're capturing. In our case, in Taxes, we capture all the all of these skills automatically. You do not need to have like a complete separate interface where users you go and create skills, that won't work. In many cases, you need to like look at the product usage and figure out whether you can create a skill for that particular use case or not. So, we automatically do in our case, and a prime example you would have used a product called this before. They also have like automatic skills. So you use and it keeps on learning as you use the product.
监控与控制:建立深层人机信任
后两个核心支柱关注于确保用户拥有足够的把控力与安全感: 3. 监控:因为代理属于后台异步运行,所以必须建立任务清单(Task List: 展现后台流式作业当前进度与处理状态的可视化列表)来跟踪每项任务的进度,并提供明晰的价值追溯(Trace: 让用户可以追溯 AI 代理在处理每个具体数值时的原始参考源与生成路径的可视化链条)。这是建立用户信任的根本。 4. 控制与干预:平台必须给用户“能随时夺回控制权”的安全感(就像握住方向盘,而不是彻底抛弃汽车)。如果 AI 代理遇到了冲突或需要做出主观假设,系统应当自动暂停,允许用户像在 Slack 等沟通软件中一样直接 @ 代理进行回复指导,解决冲突后再继续运行。此外,对于覆盖已有数据等不可逆的高风险敏感操作(Dangerous Actions: 涉及覆盖已有数据、资金流转等一旦执行便不可逆转的操作),系统必须在执行前呈现清晰的计划书(Plan: AI 在执行高风险操作前生成的包含步骤与预期结果的待审批方案),待用户审批同意后方可继续。
Original English
Now monitor. Monitor is monitoring is a key key part right. So since agents are long running now the whole point is they'll be like multiple tasks that are running and the users need to keep track of what is happening in this world. So so one simple example here could be a task list that you can do to keep track of where the processing is for each of those tasks. And second example here is like building traces in your product, you know. Trace how the agent did the work that it did. They're long running tasks so there will be multiple pieces that are moving. So you need to trace back. So in our case we trace back each and every value that the AI agent produced in a particular, you know, in the format that the users can easily see. This is where the trust is built. So it's and this is where the visibility is built so it's it's paramount that you know, you build this correctly. This is where most of the complaints can be get getting addressed if you build this right. And lastly, think of it this way, right? Control. So when something requires a judgment or when something goes wrong your platform should inspire confidence that the users can take back control. This is critical because level three product or the conveyable product that you're building will allow you the your customers to, you know, schedule a lot of tasks. But if they don't have the confidence that, you know, if something goes wrong they can take them back control then then the users will completely lose trust. So it should feel like, you know, they're taking the users are taking the wheel, not abandoning the car and, you know, creating a new car, for example. So ideally it should go like you pause the belt, fix the problem, you start it back up again. In our case we could do it very simply by, you know, pausing whenever the agent was trying to make an assumption, we pause. And then the users would come in just like in Slack or any other chat platform. They can come and tag the agent and respond with how to deal with that particular conflict. Okay, some bonus points I said. Um so, just as physical bank branches didn't disappear when the mobile banking arrived or you know, when the level two arrived, um I don't think when the agent take layer up to just it's an abstraction over the last two layers that you know, all other mobile apps and things like that will disappear. Uh the point here is a user can user will only come and delegate work they believe that you know, they can always take back control. So, so first step is you know, and having giving that the confidence that they can take back control, you know, when they when something goes wrong. So, you need to build level two features into your product as well. That's the key point here to build trust. And second thing is some actions are usually irreversible and are dangerous actions. And in those cases, you need to present a plan. So, think of it like in case of bank transactions, for example, before making a bank transaction, the plan should be there, you know, if the certain amount of value is going somewhere else. Same way in our cases, when we before we do data entry into your tax software, which can erase their existing data, we create a plan for them so that they can approve the plan before they move on. These are the key pillars that allows customers to feel you know, control and allows customers to feel as if you know, if something goes wrong, they can take back control at any point in time.
指标重构:从WAU迈向WAS
随着产品交付价值的方式从“用户亲自操作”转变为“用户委托代理”,企业衡量产品价值的核心指标也必须发生根本性重构。
- 过去 SaaS 时代的核心指标是周活跃用户数(Weekly Active Users: 传统 SaaS 产品中衡量每周有多少用户登录并使用软件的活跃指标,简称 WAU)。
- 在代理委托时代,成功的标志是周活跃会话数(Weekly Active Sessions: 由人类或 AI 代理在后台独立运行并完成的任务会话总数,简称 WAS)的上升。
我们的目标应当是让用户的周活跃用户数(WAU)下降,而周活跃会话数(WAS)上升。用户花在平台上的时间越少,越证明他们对平台的信任,越愿意将繁重的任务托付给 AI 异步处理。未来垂直 AI 产品的设计目标,应当是不断降低用户的无效在线停留时间(WAU ↓),同时处理更多复杂的后台会话(WAS ↑)。
Original English
And lastly, since we're building products now completely differently, like we're not building for the users to come in and use the product, we're building for users to come and delegate work. It's important that we change the way we measure this value. So, the most popular metric that customers that the the products use today are weekly active users. Now, this is great in the in the era where, you know, people used to come and do the work in the platform themselves. But, it doesn't really translate well with agent and delegation. So, the point is we we have to change the way we measure. And the I think the the appropriate measure here is weekly active sessions. It's I treat it as, you know, a task that is completed by a human or an agent even when the user is not in your platform. So, uh your actual aim should be the weekly active users go down while weekly active sessions go up, right? Because you want your users to have enough trust in your platform that they can come and delegate your task as much as possible. And those tasks are executed without a human help as much as possible. So, the number of sessions should go up. Weekly active sessions should go up while number of weekly active users in your platform should ideally go down. It should not be zero, of course, but should go down. So, what are the key takeaways? The key takeaways are first of all, design for delegation, not participation. Design your products in a way that users are coming in to your product to delegate work, not do it themselves. Secondly, to do to make that possible, you need to think of your product as a conveyor belt. So, users are now the supervisors, they're not operators. They're coming here to delegate work and monitor, you know, figure out if there are problems and then take actions. And finally, you need to change the way you measure how the work is done, uh not the time on the platform. Uh so, weekly active sessions instead of weekly active users. If you keep these three in mind, you will be able to build a very successful uh vertically air product, I think. All right. Uh thank you for your time. Um and good luck out there.
📌 文中提及的人物和组织
公司/组织: Filed