Abridge:AI驱动的临床智能与医疗效率革新 Latent Space 2026-05-14

Abridge的核心使命与AI医疗愿景

Abridge 作为医疗系统的临床智能层(Clinical Intelligence Layer: 整合和分析临床数据以提供洞察和支持的系统),其核心使命是利用人工智能技术,显著减轻临床医生的工作负担并提升医疗服务质量。目前,临床医生每周花费10到20小时在文档工作上,这种现象被形象地称为**“睡衣时间”(Pajama Time: 医生在下班后穿着睡衣在家中完成未完成的文档工作),加剧了全国范围内的医生短缺**(Doctor Shortage: 医疗专业人员不足以满足患者需求的情况)。我们认为,患者与临床医生之间的对话是医疗保健领域最关键的工作流程,因为无论是诊断、治疗、账单还是支付,几乎所有医疗活动都源于这些对话。Abridge最初专注于通过对话来减少医生在文档上的负担,但我们对未来的发展充满期待,目标是成为一个更广泛的临床智能层,不仅帮助医疗系统节省和赚取更多资金(Save and Make More Money: 提高运营效率和收入),最终实现挽救生命(Save Lives: 通过改进医疗决策和流程来提高患者生存率)的愿景。我们的产品理念是像空调(Air Conditioning: 在后台默默运行,只在必要时才引起注意)一样,在后台无缝运行,持续优化,只有当存在重大临床风险(Great Clinical Risk: 可能对患者健康造成严重负面影响的风险)且及时干预至关重要时,才主动介入。在这个过程中,上下文(Context: 与特定情况相关的背景信息)是理解和有效行动的关键。

Original English Source

The first and most important thing is context is everything as Tai alluded to. And I also think about how do we go from being reactive alerting to really proactive intelligence at the point at which it matters most. One thing we like to say is we want our product to feel like air conditioning. It should be in the background just making things better. And maybe if and if there is something that has great clinical risk and we're acutely aware that intervening now and not later is incredibly important, we should decide to act. Uh Bridge is a clinical intelligence layer for health systems. We really started with documentation and building for clinicians. And we think that, you know, as we think about reducing the burden that clinicians have, they're spending 10 to 20 hours a week on documentation. There's a massive doctor shortage in the country. We also think that conversations between patients and clinicians are probably the most important workflow in healthcare. It's obviously where care is given and received, but if you think about the 20% of our GDP that goes towards healthcare, almost everything is a derivative of that conversation, whether it's the claim, the payment, the actual diagnosis given the treatment. And we've started with a conversation to reduce the burden for doctors on documentation, but we're really excited about the path ahead as we become this broader clinical intelligence layer. We really think about our second and third acts around how do we help health systems save and make more money. Health systems are operating with, you know, record low operating margins. It's getting harder and harder to serve patients. And they have regulatory, some tailwinds, but also a lot of headwinds coming their way. And we think AI is ripe for helping on the saving and make more money piece. And then ultimately, how do we help save lives? The fact that our software and our product is open millions of times a week before, during, and after a patient walks in the room um gives us massive opportunity with products like clinical decision support, what Chai is building, but so many others to actually improve patient outcomes and probably one of the most important workflows and and problems to be going after right now.

AI在医疗中的挑战与机遇:高风险、垂直化与环境感知

在医疗领域应用AI,既面临独特的挑战,也蕴藏巨大机遇。我的前公司Glean(Glean: 一家专注于企业搜索和知识管理的AI公司)主要解决企业内部信息搜索问题,而Abridge则专注于医疗。两者在核心洞察上存在相似之处:上下文(Context)是驱动卓越模型表现的关键。然而,医疗领域的特殊性带来了显著差异。首先,医疗环境的高风险性(High Downside Risk: 潜在的负面后果极其严重)要求极高的严谨性。例如,开错药可能导致患者过敏甚至致命,这与企业搜索中“回答错误”的后果截然不同。这种高风险性深刻影响了我们的评估策略(Evaluation Strategy: 衡量产品或模型性能和安全性的方法),包括离线评估和渐进式发布(Progressive Roll-out: 逐步向用户发布新功能以降低风险和收集反馈的策略)。其次,Abridge的业务模式更具垂直化聚焦(Vertical Focus: 专注于特定行业或领域)而非Glean的水平化通用(Horizontal Generalization: 适用于多个行业或领域的通用解决方案)。尽管医疗领域也存在不同专科和医院系统的差异,但其变化范围相对较窄,这使得产品团队能够更集中地解决特定问题,尤其是在技术成熟且需要构建前所未有的新产品时。最后,Abridge的独特之处在于其环境感知AI(Ambient AI: 在后台持续运行,无缝收集信息并提供辅助的AI系统)模式。我们从一开始就设计产品在后台持续监听,提供无缝、主动的帮助,这与贾维斯愿景(Jarvis Vision: 钢铁侠电影中智能助手贾维斯那样无处不在、主动提供帮助的AI愿景)不谋而合。这种无缝的AI体验,让医生无需频繁查看屏幕,专注于患者,是AI的最高境界。

Original English Source

I'm Chai. I work on clinical decision support at ABridge. And so I think as Jenny said that we have this we're uniquely situated where we started off with the clinical note. What I'm really excited about and where we're expanding towards is what are all the things you can do before the conversation during the conversation and after the conversation if you did have access to all the context about patients pair guidelines medical literature and put that together and to serve you know what how healthcare could look fundamentally different. Yeah. And that's like the context engine that you guys have. Is that what it's called? Okay. Uh so historically as I understand it company started in 2018. Uh a lot of people would be familiar with like the AI voice notes form factor that that doctors would be like well do you consent to be being recorded? It replaces handwriting and what have you. Uh but I it it sounds like more recently there's been a big transition in the company or just tell me about like the the broader transition. Yeah. So from a transition perspective, we really think about our journey as how do we, you know, first chapter was first act was how do we help save time and that's where a lot of that original product was which like by the way one of the interesting stats on your landing page was like people spend doctors spend like time after hours. They call it pajama time. Okay. Why is that pajama time? Uh doctors after work in their pajamas at home or just writing and catching up on their notes every day. And you know, I think some of our favorite customer love stories. We have a Slack channel called Love Stories. We have clinicians telling us a bridge has helped us, you know, from retiring early. We're now finally able to go home and eat dinner with our kids for the first time and save their marriage and some. Yeah. One of your quotes was like we're not divorcing anymore. Like why? Like cuz they're working too much, I guess. Yeah. But um in terms of where we're going and where we're expanding, we really think about our second and third acts around how do we help health systems save and make more money. Health systems are operating with, you know, record low operating margins. It's getting harder and harder to serve patients. And they have regulatory, some tailwinds, but also a lot of headwinds coming their way. And we think AI is ripe for helping on the saving and make more money piece. And then ultimately, how do we help save lives? The fact that our software and our product is open millions of times a week before, during, and after a patient walks in the room um gives us massive opportunity with products like clinical decision support, what Chai is building, but so many others to actually improve patient outcomes and probably one of the most important workflows and and problems to be going after right now. I mean I think one thing that's that's interesting chai is obviously you came over to a bridge from glee and I think about clinical decision support uh which is you know for our listeners is basically you know in the context of a visit helping a doctor figure out the right type of care it's really a search problem in many ways right of of going through lots of different data sources very analogous to your previous role as as as one of the uh earliest engineers over at Glean um I'm sure a lot of our listeners are curious what's uh similar about the problem set you're going after now and what feels different uh now that you're you're in healthcare Yeah. Um, very similar. And I I think taking a step back, I think with every wave, there's a lot of like very similar patterns that happen across different products. A lot of social networking products look the same. A lot of like crowd-based products look the same. And I think we're seeing that's very similar in the agent era with many companies, of course, in Redpoint's portfolio and so forth. Um, and the key insight between both companies is that like you have amazing models, but like context is king and context is what actually puts them to work. Um so I see in a lot of ways a lot of similarities and like this is a healthcare coded version of clean but I think the differences are really interesting. A couple things that come to mind. First and foremost uh like the rigor at which in which in in the setting we are in um the downside risk is extremely high here in healthcare. It can actually be fatal in some cases. You prescribe something that the patient is allergic to for example. Whereas at Glean it's like oh you got the question wrong. it wasn't the end of the world in most most cases. And so what does that mean? That shapes our evaluation strategy, both offline evaluation, progressive roll out, and there's a lot more we could kind of go into there. Second thing that comes to mind is like vertical versus horizontal. Um, in both cases, there's there's a large variance, but when Glean is it's a much more horizontal company, there's a variance of personas, companies that you're working with. Um we also have a variance of uh personas, different types of specialties, different hospital systems, but the variance is a little more narrow. So from a product perspective, you're able to focus far more, especially when you have a maturing technology and you're building new products that never existed before. It lets you go specific uh go after them much more easily and especially in healthcare where so many problems have were solved with labor and process that's actually extremely ripe for AI to keep helping augment and enable. Um, and then the final thing that I think that's really interesting, Bridge specifically compared to many other companies in the AI area is the modality we started with. We're we're ambient and we're always listening in the background. And I think many more AI products will go that way, but it's actually how we started. And and I think that's actually the like the greatest form of AI we can create. AI that's actually seamless. You're not actually looking at your screen. It's all always there. It's always helping you out and being proactive. you know the Jarvis vision that like every hackathon I went to over the past decade there was always a Jarvis competitor but I actually think a bridge very much started from the opportunity and continues to go that way.

智能干预策略:从被动警报到主动辅助

在医疗领域,警报疲劳(Alert Fatigue: 由于系统发出过多或不相关的警报,导致用户对其产生忽视或麻木的现象)是一个臭名昭著的问题,超过90%的警报会被医生忽略。解决这一问题的关键在于从被动警报(Reactive Alerting: 在问题发生后才发出通知)转向主动智能(Proactive Intelligence: 在问题发生前或关键时刻提供相关信息和建议)。我们的产品目标是像空调一样,在后台默默运行,只有当存在重大临床风险(Great Clinical Risk)且及时干预至关重要时,才决定采取行动。

与其在医生与患者进行严肃敏感对话时频繁弹出警报,我们更倾向于在医生进入诊室前提供诊前准备(Pre-visit Preparation: 在患者就诊前为医生提供相关信息和建议)。例如,Abridge可以总结患者的最新病史,并根据就诊原因推荐需要讨论的事项。这样,医生在开始对话前就已做好充分准备,而非在诊疗过程中被产品多次打断。

当然,在某些情况下,实时干预是极其重要的。以预授权(Prior Authorization: 保险公司在提供某些医疗服务或药物前要求预先批准的流程)为例,患者因膝盖疼痛就诊,医生开具MRI检查。传统流程下,患者可能在几周后才被告知MRI未获批准,需要再次就诊。而在Abridge的场景中,我们可以在患者离开诊室前,悄无声息但有效地(Quietly but Effectively: 不干扰主流程但能提供关键信息)提醒医生。例如,系统会提示医生询问患者是否接受过物理治疗,以及疼痛是否持续超过六周,因为患者所投保的Etna(Etna: 美国一家大型健康保险公司)计划在加州要求满足六项条件,其中四项已通过系统上下文确认,但最后两项需要医生当场核实。如果能在患者离院前完成这些核实,MRI就能立即获得批准,从而节省了时间、金钱,并改善了患者体验。这种在关键时刻捕获医生并提供必要信息的能力,体现了临床实用性(Clinical Usefulness: 对临床实践有实际帮助和价值),实现了“节省时间、节省金钱、挽救生命”的多重目标。

实现预授权的复杂性在于需要整合海量数据。这包括与电子健康记录(Electronic Health Record, EHR: 存储患者医疗信息的数字化系统)集成,以获取患者的既往实验室检查和影像资料;同时,还需要收集并理解所有不同的支付方政策(Payer Policies: 保险公司关于医疗服务报销的规定),这些政策因州而异,且常以非结构化的50页PDF文件形式存在。因此,AI模型的高质量和与临床工作流的深度整合是成功的关键,目标是将通常需要数周或数月才能完成的流程缩短到几分钟甚至实时。

Original English Source

Yeah. And one thing I think is super interesting then from a product perspective is you have this always on seamless in the background and then you have to decide like when do you kind of break the wall almost and like say hey you know uh clinician like you might not have thought about X or or whatever it is that you want to do. Um, and obviously I think in in healthcare traditionally there's been this idea of alert fatigue and just like a million pop-ups and then a doctor just ignores all of them. It's probably a pattern that a lot of builders are thinking through now. How do you think about like the right way to uh to intervene or to pop up in in a doctor visit? Yeah, it's such a good question. I think alerts are notorious in healthcare specifically. I think over 90% of alerts are ignored. I think the first and most important thing is context is everything as Chai alluded to. And I also think about how do we go from being reactive alerting to really proactive intelligence at the point at which it matters most. One thing we like to say is we want our product to feel like air conditioning. It should be in the background just making things better. And maybe if and if there is something that has great clinical risk and we're acutely aware that intervening now and not later is incredibly important, we should decide to act. But I think if you think about proactive versus reactive, instead of alerting a clinician during a visit when they're with their patient having a pretty serious and sensitive conversation, how do we actually prep a clinician before they walk into the room with that patient? And so, historically, clinicians might have to manually go through charts with a patient that they've had over the course of months or years, and they'll try to sus out what are the things they should be doing. You can imagine a world with a bridge. will summarize all of the most recent contacts for you, tell you based on the reason for a visit the patient is coming in for the types of things you should be discussing. And so you're actually going into that conversation prepped rather than walking in cold to that patient visit and then having this product interrupt you five or 10 times throughout the the visit. And there might actually be times where it's really important to interrupt. We have a product called uh prior authorization. And so this is when you may go into a doctor's office with knee pain. They'll prescribe you an MRI and I think so many of us have had this experience before where in 4 weeks you'll get a call saying, "Hey Sean, that MRI that you are prescribed wasn't approved and why don't you come back in? We'll figure it out." In a world with a bridge, we might choose to actually quietly but um still alert a doctor in that visit. And alert is probably not even the word we would want to use. Before a patient leaves, we would want to tell the doctor, "Hey doctor, before Sean leaves, you should ask him, has he had physical therapy and has his pain lasted for more than 6 weeks?" Because the Etna plan that he's on in California requires six things. We've already confirmed four of them have been met because we have all the context. But these two last criteria, if you can address with Shawn before he leaves the room, we could actually guarantee that your MRI is approved before you leave. And so when you think about clinical usefulness, um impact to the patient, I think there are instances in which if we can catch a doctor while the patient is still in the room, um you know, as we think about save time, save money, save lives, you kind of get to check all of those boxes. But um when you know doctors have 15 minutes between visits, we have to be really really thoughtful about when it actually matters. I think there's this interesting product opportunity that AI can have is reduce latency in the world. Um for example, prior authorization is an example of where care gets delayed and so great AI can reduce that. And I think the problem with alerts before partially is a technical problem like it's the quality of your alerts really matters. they're going to get ignored if you get alerts that similarly in engineering where they're noisy alerts that you can't act on. But if you can make really high quality alerts with with both the context as Janie said and really high quality models, then I think you can create a whole another game. Yeah. And I really like that experience because I think it kind of starts to tease apart like what makes this so hard and unique. I think like one to make that prior authorization example possible, think about all the data that you need to have. You need to integrate within with the electronic health record to know all of the patient context. Do we have access to your previous labs, previous imaging, and then to actually match you and to know that you're on Etna? We have to collect all of the different payer policies and they vary by state. Some of these payer policies live on websites, some of them live in unstructured 50page PDF files. I thought this episode was to make sure we didn't scare people from healthcare. But uh but it's I think when you think about the things that make it hard, it also gives you the moat. And then I think the second is the uh AI and the model quality we need to be able to hang our hat on. And so the bar I think similarly when I worked at opendoor I worked on pricing models like every outlier wiped out the margins of 30. And so similarly here in healthcare the the bar for accuracy is so high. And then I'd say the last is workflow is everything. You know, if insurance companies deploy AI, it typically happens too late. And this is when you have the notorious kind of like comical examples of AI just fighting each other when it's too late. But if we can pull forward the use of both the AI, but also the ability to solve problems when the patients in the room, you can start to collapse what typically takes weeks or months after your visit. um ideally down to to minutes or real time. And I think um it's it's where healthcare is both very difficult but also extremely rewarding if you can crack it.

产品形态、个性化与多方生态系统

Abridge的环境AI(Ambient AI)产品形态主要通过移动端(Mobile: 手机或平板电脑等便携设备)和桌面端(Desktop: 个人电脑或工作站)提供。临床医生通常在诊室之间使用移动设备,而在完成笔记或准备第二天工作时则使用桌面设备。我们正积极探索与室内设备(In-room Devices: 部署在诊室内的智能设备)进行合作,以利用多模态数据捕获更多未被记录的上下文信息。此外,增强现实眼镜(AR Glasses: 提供叠加数字信息于现实世界视图的眼镜)也是一个未来方向,它能在不依赖屏幕的情况下,将实时信息直接呈现给临床医生,让他们专注于患者。

个性化(Personalization: 根据用户特定需求和偏好定制产品或服务)是Abridge产品策略的核心,我们从三个层面实现:

  1. 个体医生层面:医生的诊疗笔记是其工作和护理方式的深刻体现。因此,产品需支持医生对风格(如项目符号或段落、简洁或全面)、常用短语和笔记模板的偏好。我们持续收集用户反馈,确保这些风格偏好(Stylistic Preferences: 用户在内容呈现形式上的个人喜好)不会影响准确性(Accuracy: 信息或结果的正确程度)和质量(Quality: 产品或服务的优劣程度)。
  2. 专科层面:不同专科的医生工作流和文档要求差异巨大。例如,心脏病专科医生与皮肤病专科医生的笔记和工作流截然不同。Abridge需要针对每个专科进行深度定制和校准,以确保生成的笔记既完整又符合合规性(Compliance: 遵守相关法规、标准和政策)和可计费性(Billability: 医疗服务能够被保险公司或患者支付的能力)。
  3. 医疗系统层面:大型医疗系统通常拥有多年甚至数十年积累的最佳实践(Best Practices: 在特定领域被证明最有效或最成功的做法)和临床指南(Clinical Guidelines: 为特定临床情况提供建议和指导的系统性陈述)。Abridge的临床决策支持产品能够整合这些医院内部指南,在诊前、诊中或诊后为临床医生提供具体指导。这种深度整合不仅增强了产品的护城河(Moat: 公司相对于竞争对手的竞争优势),也使Abridge成为医疗系统值得信赖的合作伙伴。

Abridge的成功也得益于其独特的数据飞轮(Data Flywheel: 通过数据积累和反馈循环不断优化产品和服务的机制)。用户在产品上的每一次编辑和反馈,都成为宝贵的训练数据,推动产品不断深化个性化。医疗保健是一个多方利益相关者的复杂生态系统,包括医生、患者、支付方(Payers: 如保险公司)和药企(Pharma Companies)。Abridge致力于将这些原本独立、昂贵且复杂的系统整合到一个统一的平台中,从而提高效率,并为所有相关方带来更好的结果。

电子健康记录(EHR)系统的深度合作是Abridge成功的基石。我们必须与所有合作的EHR系统保持极其紧密的伙伴关系,确保数据的双向流动。对于临床医生而言,他们大部分时间都在EHR系统中度过,因此任何新增的产品都必须减少点击(Save Clicks: 减少用户在操作界面上的交互次数),而不是增加负担。Abridge通过与大型EHR厂商的紧密合作,利用非标准API实现数据拉取和推送,确保产品能够深度集成(Deeply Integrated: 与现有系统紧密结合,无缝运行),从而获得真实的用户使用和采纳。我们坚信,除非产品具有互操作性(Interoperability: 不同系统之间交换和使用信息的能力),否则它将无法被广泛使用。Abridge将自己定位为构建在EHR之上的临床智能层(Clinical Intelligence Layer),连接提供方、药企和支付方,开创一个全新的医疗智能格局。

Original English Source

Just to get some baseline on the form factors because I've seen some videos on on your website and stuff. You guys talk a lot about ambient um AI. Uh is it primarily on the phone? Is there any other form factor that people get a bridge in? Like is there like a bridge room setup where it's just like always on like I don't know a bridge podcast studio. Yeah. Primary form factor is mobile and desktop. Usually clinicians are walking in and out of rooms with mobile but at the end of the day when they're closing out their notes or wanting to prep for the day ahead they might use desktop. We have been having a lot of um really interesting partnership conversations with a lot of these in room device companies. as you think about what is, you know, the power of multimodality and even more data as you think about all of the what is today not captured context. Um really really fascinating to think about especially even as we go into uh building and scaling our nursing product. It's one where nurses constantly, you know, as they're walking in to check in on a patient for 2 minutes or maybe even 30 seconds. starting on a bridge uh experience is probably going to take longer than the visit. And so what can we do with in room devices that are always on starts to beg really really interesting and and fun questions like the way in in tech companies we have all these Google meet things we might as well set up entire rooms with just a bridge tech very much and and I also think similarly about like actually also AR glasses and so forth I think is also quite relevant where part of it is how do we bring it in a way without like a screen but like also bring the information to the clinician in real time but also to let them focus on the patient. Do you think they want that? I'm I'm just like very I'm tend to be skeptical AR, but you know, I'm curious what you've tried. You know, admittedly, it's not a near-term product road map by any means, and I'm here being such sick AR stuff for surgeries actually when people are trying to visualize like the, you know, you're about to make an incision, but you want to see like what the cut might look or what the body might look like inside, and they can basically layer in imaging. Uh yeah, that's cool. Yeah. Yeah. Yeah. some point in the future but there are a lot of uh our largest customers and and at the largest health systems integrating already and so even as we think about building into it I think unlocks a lot of product capabilities yeah and just to establish the terminology sorry and I know I'm like I'm asking basic questions somewhat for myself but also for the audience who might be less integrated when you say health systems it's like the Johns Hopkins the Kaiser permanent these are your customers right and the outcome that you deliver for them is happier doctors ers uh reduce like whatever cost like I guess of like of processing reduce mistakes. I think it's it's it's weird in a sense that I feel like there's also like a secondary customer like the customer of the customer and I don't know if you do you think about it that way? Yeah, we have I think the other interesting and complex part of uh building product is we have our buyers who are the chief medical information officers, the chief financial officers, the CIOS of these large health systems. Our users today are clinicians. But if you think about who downstream is impacted, it's patients. And so as we build um with every product in mind, we think about who are we building for? Who's the secondary user? And what does that mean either in terms of experience, security, compliance, uh ROI that we have to make tangible and so like you said like time savings is one of them. But for CFOs, they care a lot more than just time savings. we have to show for every dollar you put into a bridge because you have more compliant documentation or because you have fewer queries coming from your billing team um we actually save or add real dollars um to your bottom line or topline I think are things that we're constantly thinking about because of the the dynamic across all three sets of users and I think there's a whole another axis too with the with the payers and pharma as well and like connecting all these three big stakeholders and healthcareers do the payers ever see your data. Sorry, the payers mean the insurers, right? Yes. They they also see a bridge data. Uh direct um they wouldn't see the the raw bridge data, but where when you're working together on something like prior authorization, whatever information they need, we we'd communicate to them. All right. Yeah, that's cool. I guess you know would love to dig into just obviously you have uh you guys solve a lot of problems on the AI side and so maybe to start at the highest level what's one of the hardest problems you have to solve in AI at a bridge today Yeah, it's such a good question. Personalization is massive for us. We think about personalization at three levels. The first is at the individual, the second is at the specialty level and then the third is at the health system or the organization level. To your point, there are a lot of individual preferences. Um, when a note is produced, it almost is a reflection that is so deeply personal of a doctor's work and how they give care. And so, do they have preferences on things like style? They might want bullets versus paragraphs, really concise versus comprehensive. Um, they also might have phrases that they really like to use or the templates that they want every note to be structured. And I think, you know, we see it in our feedback all the time. We want two spaces in between sentences or I refuse to use this tool. And so that's something that we've had to build in. And I think the tricky part is how do you make sure that stylistic preferences don't actually interrupt accuracy and quality. And that's something that we've really had to refine and hone over time. Second is at the specialty level. A cardiologist note or workflow is going to look very very different from a dermatologist workflow. I assume cardiology notes are the highest stakes for you guys given your CEO as a cardiologist. just like our little guy. Sh our CEO is still a practicing cardiologist. He rounds once a month and so uh first call when we want just quick and easy user feedback too. But uh specialties require a lot of personalization both in terms of what does the product actually look like and so we make sure that as new users on board we catch that and the product kind of proportionally reflects that but also on the back end like eval at the specialty level they are hardearned to actually calibrate and get right. What does a really great dermatology note look like? What actually makes it complete? what makes it compliant and billable is very different than a primary care doctor. And so it's not just about what does the product experience look like, but on the back end tuning and really really deepening our understanding for the specialists. What does great output look like? And that's um obviously a problem that we need to calibrate internally, externally, online, offline, but um takes lots of cycles but is necessary in a high stakes environment. And then at the health system level um for products like clinical decision support you have health systems who've spent years or decades refining their best practices and they want to know hey we love your clinical decision support product but how do we embed our own hospital guidelines into them to actually inform clinicians before during or after a visit what breast best practices should look like. And I think as you think about like deepening moes as well when health systems uh trust us with that data allow us to kind of productize it and directly into the clinical workflow uh I think makes us a really really great partner to health systems who want to build something that truly meets their needs their practicing guidelines and I want to add on to that uh the for the clinical documentation problem it's it's very similar to like AI writing writing that doesn't feel like your own and then we call that slop but the way I describe one framing of slop is like AI without context but we have all that context you know and both the clinicians uh can have it and can guide it and so part of the other interesting exhaust for us is like memory is like actually one of these new systems of records almost and we also have all the edits people make on our product and when you think about a data flywheel and how we get better over time becomes really really powerful as a mechanism to just going deeper in personalization It's so interesting. I love this idea of like working with systems on, you know, the guidelines they built up over a long time. I feel like so many of the best AI app companies today are, you know, the question is how do you take, you know, the expertise that a law firm or or a bank has built up over many years and then add that as context and also a special sauce over like a an AI tool. And so seems like you all are really doing that very effectively. Yeah, we're now starting to have our customers ask like what are other customers doing and how are they doing it? And I think as we think about having visibility across such a large set of care being delivered right now, a really interesting place we could also partner.

AI技术深层挑战:实时性、成本与专有数据优势

Abridge在AI领域面临的核心挑战是如何在医患对话中实时(Real-time: 在极短时间内响应和处理信息)提供智能辅助,同时平衡质量(Quality: 模型输出的准确性和实用性)、延迟(Latency: 系统响应时间)和成本(Cost: 运行模型和基础设施的费用)。在任何AI产品中,这三个关键绩效指标(Key Performance Indicators, KPIs: 衡量业务或项目成功的关键指标)都是需要权衡的。我们希望在对话中实时指导临床医生,但又不能因此大幅增加成本。这需要非常智能的模型,因为我们处理的是大量交叉数据和复杂的上下文信息。为了避免警报疲劳(Alert Fatigue),我们必须确保高智能和高质量的输出,同时保持快速和成本效益。这需要大量巧妙的工程设计,例如将政策建模为某种中间表示(Intermediate Representation: 一种抽象的数据结构,用于在不同系统或阶段之间传递信息),以使问题更易于处理。

在选择AI基础设施和模型时,我们采取了深思熟虑的策略,以应对不断变化的AI格局。我们关注第三方模型(Third-party Models: 由外部供应商开发和提供的AI模型)的发展趋势,并识别我们能发挥独特优势的领域。虽然通用模型会持续改进,但我们认为专有模型(Proprietary Model: 由公司内部开发和拥有,不对外公开的AI模型)在特定场景下能带来显著优势,即提供更高的质量或在相似质量下实现更低的成本和延迟。这种优势来源于我们独特的专有数据(Proprietary Data: 公司独有且不对外公开的数据集)。我们拥有数亿次的医疗对话数据,这是一个非常独特的数据集(Data Set: 用于训练和评估AI模型的数据集合)。这些对话数据构成了患者与提供者之间互动的代理追踪(Agent Trace: 记录AI代理在执行任务过程中的每一步操作和决策),这正是医疗领域“调试”发生的地方。通过这些大规模的追踪数据,我们能够训练出在特定用例(如转录、说话人分离(Diarization: 识别对话中不同说话人的技术)和笔记生成)上表现更优、成本更低、速度更快的代理。

我们深知模型提供商正在更多地关注代理工作流(Agentic Workflows: AI系统能够自主规划、执行多步骤任务并与环境交互的工作流程)的训练。同时,由于医疗查询是消费级模型(Consumer Models: 面向普通消费者设计的AI模型)提供商的重要应用场景,他们可能会优化模型以在权重中编码大量医疗知识。这对于Abridge而言是利好,因为开箱即用模型(Off-the-shelf Models: 无需定制即可直接使用的通用模型)在通用医疗信息处理方面会持续改进。我们的策略是采用模型组合(Constellation of Models: 结合多种不同AI模型以实现特定目标的策略),根据具体需求选择最合适的模型,最终目标是提供最佳的产品体验。

随着模型能力的提升,我们看到了新的可能性。例如,我们发现几乎所有AI代理本质上都是编码代理(Coding Agent: 能够理解、生成和执行代码的AI代理)。在医疗领域,电子健康记录(EHR)可以被视为一个庞大的文件系统(File System: 组织和存储计算机文件的方式),其中包含海量信息。如果AI代理能够更好地操作和读取这些数据,将其视为文件系统进行处理,将极大地促进我们所有产品用例的实现。

关于实时性,目前我们的系统主要基于批处理(Batch Basis: 集中处理大量数据而非实时处理)模式。然而,我们正在开发有趣的原型,探索如何在对话中根据特定时机触发模型和代理工作流,以降低反馈循环延迟(Feedback Loop Latency: 从系统接收输入到提供反馈所需的时间)。虽然我们尚未实现完全的语音输入/文本输出(Voice-in/Text-out: 接受语音输入并生成文本输出)或语音输入/语音输出(Voice-in/Voice-out: 接受语音输入并生成语音输出)的实时交互,但我们正通过巧妙的工程设计来降低延迟。目前,我们主要采用文本输出,因为在医患对话中引入第三个“声音”可能会过于干扰(Disruptive: 造成中断或不便),影响医生专注于患者。

Original English Source

to make things simple let's take like building off the prior o example so one thing Jamie talked about is like okay this data is all over the place and there's this cominatorial explosion of like uh procedures uh payer policies and even sometimes different health systems there can be some cross productduct of all of these different considers ations you have to take in account. But what's really really really hard about this problem is actually doing it real time in the conversation. So you know in any AI product usually the three KPIs you care about are quality, uh latency and cost. Now what we're saying is we want you to do this real time in the conversation guiding the clinician. How do we do it in a way that does not break the bank? Um but we're using but we also need very intelligent models because you're working with this cross productduct of data and this like all this context layer as well. So you need high intelligence and high quality because you don't want the alert fatigue but you also need to be fast and cost effective and so that's where I think a lot of clever engineering goes actually. It's like okay without getting into all the details here can you model these policies in some intermediate representation or other things that you can do that can actually make this problem tractable? Um and of course the parado frontier is always changing but we're also trying to do this now. Yeah. What implications has that had for what you take off the shelf um and say you know what we don't need to be world class at X we'll just take this from the model providers or from some infrastructure player and what you're like no this is where we spend most of our time focused on. Yeah. Um this is the the fun challenge in AI right of course with with the shifting landscape. We try to be extremely thoughtful on predicting the trends of where third-party models are going and where we can uniquely go. Um, and you know, sometimes I feel like when when you talk about AI models, we're like the models are just going to get infinitely better. But I don't think maybe in the grandness of time you could say that, but actually within every month, every quarter, there's specific ways they're getting better. You know, they're training on a lot more uh coding data to be better coding agents, for example. Um and so we have to think about where are the things that unique data that we're uniquely training on. Um or actually to step back a little like where is a proprietary model bring an advantage to us is if it can give higher quality or lower cost and latency for similar quality very similar to many other companies. Um and when we can do that is when we have proprietary data. So for example, we have on the order of 80 million or hundreds of millions actually now getting close to of medical conversation. This is a unique data set and this data set it's it's very interesting because this data set is effectively a large part of the trace between the patient and the provider. That's where the quoteunquote debugging happens in healthcare. We actually have these traces at scale as in like as uh RCO's even called it an exhaust that comes out of our product. And so when you have these traces, that's how you can actually train better agents on certain use cases um whether it's your transcription diorization use cases or or so on um or like note generation models and we can do that much cheaper and faster. But we're always also working with these third party model providers. You know, we we closely collaborate with them and that's how we kind of predict where the trends are go. The thing that I think about a lot is that um I know that the model providers are going to train much more on like agentic workflows and so forth. So that's great. So that you have a better agentic hardness. But the other thing that's interesting is you know that the model providers because a large class of the consumer model providers is healthcare queries. You know that they actually might uh optimize to train a lot of healthcare data to actually encode the knowledge in its weights. And I think this is just a great thing for us as well where the off-the-shelf models can keep getting better at general healthcare information such that what our strategy is we have a constellation of models. we can use something for this that and like we only care about at the end of the day the best product experience. Yeah. And obviously you have like overall capabilities improving. I'm curious like as as these models get better. Is there something you look at and you're like you know 3 months ago we really couldn't do that but god like the you know the latest the latest models really allow us to do it. So here's something interesting that I've kind of been toying with. Um so all models are this wasn't super super obvious a year ago but now it's become clear and clear that almost every agent is a coding agent underneath underneath the hood right so you you give it whatever a file system it can write its own code and so forth so when you think about within within healthcare and the use case that we have you can think of the EHR effectively like a file system um it's just it's a storage of all this information it's actually a lot of information there it cannot fit into the context window at least of today's models And you want to use that context effectively for all these product use cases we're talking about. And so if you have better agents that can actually like manipulate data, read that data, treat it as a file system as we see they're going and we know model companies are investing this way, um then that actually very directly benefits us. Yeah. Yeah. Okay, cool. Again, just establishing basic things, but we're going back to the model stuff. Um I I'm really interested in double clicking more on like the real-time uh element, which is pretty important for for both of you. Is it is real time basically just batches of like every 1 minute every 5 minutes? Is that how we actually do it or is there some more native like genuinely real time in the sense that OpenAI has a real-time API or or Gemini has a real time API? Yeah. Yeah. Yeah. So today it is more on the on the batch basis but there's interesting prototypes that we have that were still not fully uh full-time you know voice in text out or um in in that sense but actually can you can you trigger your models your agents or agentic workflows depending on actually the right times in the conversation. Um and so you can imagine you know different techniques to bring this latency down and like you know you want to bring the feedback loop down as much as you can. Um uh and so a lot of clever engineering there without fully maybe one day we'll do full voice in and text out and train a model to do something like that. Do people don't want voice in voice out. Right now we aren't creating experiences that are like kind of during the conversation kind of inter it's almost like might be too disruptive too too disruptive until like who knows maybe eventually you could have full voice agents once we the quality and we improve the comfort of the technology um but right now I think gra that change is much more gradual and it's more text focused text out and I think so much of currently what our product is trying to do is allow a clinician to focus on their patient and I think maybe at some point, but I think right now patients, clinicians don't want a third voice, at least in a literal voice in that room. And so, how do we be there with all the contacts and information ready at hand when there's the right moment?

评估策略、隐私合规与医疗AI的未来

在医疗AI领域,评估策略(Evaluation Strategy)至关重要。从产品开发的第一天起,我们就明确“好”的标准,其中临床安全性(Clinical Safety: 确保医疗产品或服务不会对患者造成伤害)是底线。为了确保高质量输出,我们采取多层次评估:首先,内部临床医生(Clinicians: 从事临床医疗工作的专业人员)会执行LFD(Look at the Fing Data)流程,对产品输出进行初步判断。其次,我们创建了大语言模型评判器*(LLM Judges: 使用大型语言模型作为评估工具,根据预设标准对其他模型输出进行评分),并结合标注数据和内部/外部评估人员进行校准。在发布任何重大变更之前,我们会进行全面的评估。我们的目标是将评估周期从数月缩短到数周再到数天,这既是一个机器学习问题,也涉及大量的运营优化,需要深厚的领域专业知识(Domain Expertise: 在特定领域内积累的深入知识和技能)。

隐私与HIPAA合规(HIPAA Compliance: 遵守美国《健康保险流通与责任法案》,保护患者健康信息隐私)是医疗AI不可逾越的红线。所有用于在线评估和学习的真实世界数据都必须经过去识别化(De-identified: 移除个人身份信息,使其无法追溯到特定个体)处理。我们甚至开发了专门的模型来识别和移除临床转录本中的受保护健康信息(Protected Health Information, PHI: 法律规定受保护的个人健康信息)标识符。在确保去识别化模型准确可靠的前提下,这些数据才能用于训练和评估。此外,我们与客户签订严格的数据合同(Data Contracting: 规定数据使用、存储和访问权限的法律协议),明确谁可以访问PHI数据、数据保留期限以及去识别化前的处理方式,以确保始终尊重客户数据和隐私。未来,患者可能更愿意分享其去识别化数据,以帮助其他患者从其经验中学习,这将是AI技术在医疗领域的一大突破。

Abridge的规模化运营(Scaling Operations: 扩大产品或服务的覆盖范围和处理能力)也带来了独特的挑战。在处理数亿次对话数据时,基础设施的可靠性、模型运行成本(尤其是Token使用量(Token Usage: AI模型处理文本时消耗的最小语义单元数量))以及计算效率的优化变得至关重要。我们通过后训练(Post-training: 在模型初始训练后进行的进一步优化)和效率最大化来应对这些挑战,特别是在质量提升空间有限的领域。我们认为,Abridge在某种程度上走在了未来,因为大多数AI用例仍处于用例发现模式(Use Case Discovery Mode: 探索AI技术潜在应用场景的阶段),而我们已经进入了优化模式(Optimization Mode: 专注于提高效率、降低成本和提升性能的阶段)。

我们对医疗AI的未来充满期待,它将围绕三重目标(Triple Aim: 医疗保健领域提高护理质量、改善患者体验和控制成本的综合目标)展开:提高护理质量、降低护理延迟、减少成本。我们设想,当实验室结果更新时,后台代理能立即启动,利用所有上下文信息,建议临床医生下一步行动,并通知他们,从而显著降低护理延迟(Latency to Care: 患者从需要医疗服务到实际获得服务之间的时间)。更长远来看,AI甚至可以直接连接患者和消费者。

在团队建设方面,Abridge采取了独特的方式。我们的团队中包含许多临床科学家(Clinician Scientist: 兼具临床医学背景和技术开发能力的专业人员),他们通常是拥有医学博士学位并具备全栈工程师或高效提示工程师能力的专业人士。这些“变种人”的加入,极大地提升了我们产品的临床实用性和评估质量。他们深度参与整个评估流程,确保我们设定的评估标准是临床相关的。

在产品开发方法论上,我们从过去“快速行动、快速发布”的经验中吸取教训。在医疗这种高风险环境中,盲目追求原型并不可取。我们强调书面清晰度(Written Clarity: 清晰、准确地表达产品需求和设计思路)和战略性地选择问题。在投入大量资源开发新功能之前,我们会深入思考“为什么我们公司应该解决这个问题?”、“如果竞争对手也开发了类似功能,我们的竞争优势(Right to Win: 公司在市场竞争中获胜的独特优势)是什么?”。虽然原型在快速探索不同解决方案时仍有价值,但最终的产品需要详细的实现细节、安全合规性、边缘案例处理等,这些都必须以书面形式明确记录。

对于AI基础设施,我们认为一些为人类协作而构建的技术,如事件驱动系统(Event-driven Systems: 基于事件的发生来触发和协调操作的系统,如Kafka、Temporal)和无冲突复制数据类型(Conflict-free Replicated Data Types, CRDT: 允许多个用户同时编辑共享数据而不会产生冲突的数据结构),在未来大规模运行的代理系统中仍将具有持久价值。Abridge的未来将更加代理化(Agentic: AI系统能够自主执行复杂任务并与环境交互),后台代理将能主动响应实验室结果、连接患者,进一步扩展AI能力,最终实现一个更智能、更高效的医疗未来。

Original English Source

what about evals how in the world do you like this is such a complex product surface area um would love to hear you riff on that and also how's that evolved like I'm sure you've gotten better at it so any any kind of learnings along the way from an eval perspective uh we I think from day one when we build any new product or feature we think about like what does good look like and there are table stakes things like clinical safety, but then you start to get deeper into what does good quality look like. And when you go into something like our core product, there's stuff like style and completeness and there's things like this this note actually becomes something that can be billable which is obviously very very high stakes for a health system. We have a number of ways in which we get confidence for this. We have uh internal in-house clinicians who do what we call an LFD process to give us our very first pass at is this or isn't this a good enough output? Uh look at the effing data that's why I was smiling. I was like is JD going to mention what it stands for? That does not there's like a million acronyms I'm like supposed to know that I don't. Oh yeah, of course. Like an LFD. I've never heard of LD. It's a Birch specific version. I think I got through three days and then I had to ask someone. I thought it was just me that didn't know, but it's our I look at the data as a meme in ML because you you tend to not look at it. You just want to look at number go up. Exactly. Um but yeah but so uh we make sure we look at the data and then as we think about all of the components of good output we one create LLM judges across all of these and we make sure with annotated data and either internal or external evaluators we feel like these judges are calibrated and then depending on the stakes we also work with in-house and third-party evaluators across all of these before we ship any big change and I think the goal is in terms of evolution. How do you go from this process taking months down to weeks down to days? Some of it is like a true science and ML problem. A lot of it's also just like hard operational work. Have you planned ahead in terms of what you need? Have you really optimized the capacity that you need across all of the different specialties you need? Have you gotten a really good sense of like which third parties are great to work with for what use cases? I think this takes a lot of domain like expertise and like to be frank lots of mistakes and errors and figuring that out. And so I think as much of it is an ML problem like so much of it has also been operational gains that I think are hugely hugely important where domain specific expertise is is everything. Yeah. But it's funny because I feel like people talk about healthcare like it's one giant market and the reality is it's like you know dozens and dozens of of submarkets and so it feels like in your eval you obviously have to you know uh build that up across the board. Totally. And is specialization the the primary cardality that's the word that comes to mind sometimes depending on the product or the use case. And so if we're making a note improvement or feature for a particular specialty definitely, but we have products that are for nurses. We have products that um are really really aimed at making the document or the output a lot more billable. And so we'll actually want to work with coding teams and not necessary clinicians. And so meaning healthcare coding I see. Yeah. But is this output proportional to the work that was actually delivered? Is there sufficient documentation to justify the amount that a health system may end up charging? And so, um, specialty sometimes, but also domain very different across all of the different products that we're working for. And building out that network is, uh, not easy. And I think is is where a lot of our operational investments have gone into. And I think um I I view a lot of analogies to self-driving cars here where like part of it is we actually really want progressive roll out of of features to actually test in the real world. Is this useful? Is this going to work? You know, one one big difference compare compared to past lives is before I'd build a product maybe at alpha and then I'd like G8 the next week cuz I'm like go move fast ship and whatnot. But the mentality is is like you I want to make contact with reality as quickly as possible but I want a progressive roll out because as much as I get as large of an offline eval set I want the distribution of that to actually match real life distribution and over time by rolling out early I actually think similar to Whimo has a tagline the world's most experienced driver I actually think another thing that can actually like uh at least linearly increase for us is like both the size of our evaluation offline and online. Um that and it all feeds back. I think some thing that's been earned over time, speaking of evolution, is just the trust we've gotten with customers. Um, historically, a lot of these health systems when they bring on new vendors, their release cycles are quarters, sometimes twice a year. We've gotten our customers onto monthly release cycles, which is pretty fast for health systems. But I think what is more exciting over the last call it few quarters has been um a subset of our customers have said we actually want to innovate with you. We trust you and we have a pretty like decent chunk of our customers who say we'll actually develop with you outside of these monthly release cycles. We have a higher tolerance. We know that the stakes are very high but we want to be the first ones using these products giving you feedback. And so for a pretty substantial set of our customers, we've been able to convince them to to be able to ship, you know, in this gradual way way before G. I think something we talk about a lot internally is uh trust is earned in drops earned in buckets. And so we still can't do what I used to do when I worked at Loom. We had 30 million users. I'd just be, you know, rolling out experiments left and right. Like the bar is still quite high for iterative rollout. But I think like because of the the trust we've earned, we're able to learn at pretty high volume very quickly. Yeah. I mean your scale is still pretty pretty huge. One thing I want we were going to go into scale right in a sec. Uh one thing I wanted to follow follow up my eval again just coming from a generalist engineer point of view just thinking through what would people be scared of in doing this. Uh the privacy and HIPPA elements of this uh I have zero experience in that what do you have to do what is actually surprisingly not that bad. So one thing that's really important here from a compliance perspective is very much that any of the data we use needs to be deidentified any real world data we use as a basis of um online eval sets or learning from and so you have to and there's like actually very clear you know government guidelines what what counts as PHI and so we've actually even have built models that can take for example a clinical transcript and actually remove all the key PHI indicators and so you have a scrub/ deidentified version and then once you and so one thing that's important is first you got to get confidence in that model in the first place right and prove that out because you know now you've like multiple probabilistic systems on top of each other but once you once you have that then you can actually train on it use it for evaluation so forth provided one of the cool things also that you can do from a business side is the right data contracting as well with your partners is the anonymization one way like once it's done you cannot undo it or is there someone who holds the master key that can yeah okay so it's it's one Yeah. Yeah. I guess I mean I guess that's how it works. I just wanted to because like you know there there's a lot of these like learning from feedback and everything that like you would want to debug more but you can't because you just physically don't allow yourself to. Yeah. Right. Well, some of it's also written in our customer contracts in terms of who can or can't access PHI data. How long do we retain it um before it gets deidentified? And so we have a pretty high bar for who can access that PHI data. Um just to make sure that we always respect our customer data and privacy, but that's something that we partner with our customers on too to make sure that as we want full as close to precision as possible in in that quality, we can still use it. Yeah. But it'll be fascinating to see how that space evolves, right? is you think about uh I I used to work at a company that like did a lot of healthcare data in the cancer space and if you ask like the average cancer patient like hey do you want people do you want other patients to be able to learn from your experience like they're like please like I love like nothing more than for other people to be able to learn from the experience that I had and so obviously in the past it was a lot harder to do that kind of learning but I think with this technology uh that actually might really be practical and so it'll be fascinating to see how that uh how that continues to evolve. Yeah, there's so much in our data set of a 100red million conversations. You you can imagine things like insights that you can give to the clinician. How could you oh, how could you have reacted to this and coaching or insights around like uh which treatments are effective or like because you have this again this data source that was never captured before but that's like where like intuition or experience is created from going back to this idea that the conversation is the agent trace. Yeah. I guess you know back to the 100 million conversations. I mean, I feel like you have this insane scale that maybe only a few other AI app companies have and everyone else dreams of. So, not not everyone has had to confront this yet, but maybe just talk about some of the like challenges of of operating at that scale and what what do our listeners have to look forward to if they ever get to uh to this level of scale. I think at large and larger in scale so of course there's a general like infrastructure reliability like when in any given startup you're kind of building the plane while while it's flying. So, so there there's some notion of that. Um but I think what gets interesting on the AI and ML side for sure is this is as you get at more and more scale. So one you you have the data to first and foremost do this but like you actually start thinking about costs or infrastructure in a whole different way at scale versus like a prototype. You can use the most expensive model. You can burn as many tokens as you want. But when you're doing 100 million conversations yeah token max and leaderboards are less exciting in that context right when you're doing that. And so that comes for we have the data and we also have the team that's able to actually like post train based on this and you can actually optimize for efficiency especially in areas where you believe that maybe a lot of the quality headroom is less so um and you don't expect the other off-the-shelf models to get that wave such that you want to do you know efficiency maximization uh in terms of uh compute and tokens. Yeah. Yeah, I mean I feel like you guys live in the future in some way where most most use cases today are really just in use case discovery mode where it's like god I really hope I can find something that can get to scale and so you're always going to use the most powerful model and then the few things that do get to this level of scale you start to do those kind of optimizations. Yeah, it's a natural trajectory where it's like 0 to one we're not talking about any of these optimizations but when when maybe we're in the one to 100 or so forth then we're in optimization mode and like what works out really well is you got all this data from zero to one that lets you do this. Yeah. So that's fascinating. I mean, I feel like there's there's, you know, one thing that's so interesting about the bridge footprint is like you're in the doctor patient vit in real time. There's probably I always say there's like probably 50 years worth of product you could build on top of that. What gets each of you like I don't know what are you most excited about building, you know, either in the short term or medium-term or even, you know, long down the line? I think something that I get really excited about is that the same conversation can serve so many stakeholders. Uh if you think about the conversation, a doctor needs to know what is the documentation, how do I make sure that this fully represent the care I gave. A patient needs to know like what the heck just happened? This was really overwhelming. What are my next steps? A payer needs to know, was this the proper and appropriate care given? um a pharma company might want to know why isn't this drug being properly used or is there actually a good candidate for this clinical trial that I'm about to run and I think where I get excited is that our product and our platform and our infrastructure can be the same product across all of those things and start to what's today like separate very expensive complex systems that serve each one of these stakeholders in very different ways start to kind of collapse all of that into a singular platform that enables not just more efficiency across the board, but also better outcomes for for everyone. And I think, you know, I think all of us experience healthcare in probably very painful ways. And I think knowing that there is a world in which we can simplify a lot is is really exciting to me. And it all starts at the conversation. You know, it's interesting. I think of it very similar to going back to the KPIs that any AI product cares about. How do you increase quality of care? How do you reduce latency to care? And how do you reduce costs which is huge uh in healthare and they call it the triple aim in healthcare, right? Um but but but but very similarly to to building AI products and the thing that really excites me is when we talk about that latency piece. We talked about one example earlier of like prior authorization. Can you reduce the latency to care? But you can imagine so much more like oh as soon as the lab value gets updated do you have like a background agent that like kicks off and uses all the context to be like oh hey actually the patient should do this next for example and flagging that to the clinician who's always in the loop but reducing that latency um to care and then you can imagine this is much further down the road but it's like even connecting that to the direct patient and the consumer and so how can you how can you build build a bridge to all of these things very cool I think like the connections piece is is just an ever growing thing. Um and one one of the key partners is the EHR and I I wonder what that relationship is like like um will will they will they like you know look at this as like something that is valuable enough that they want to own someday. Yeah, I think our partnerships with EHRs like we know that we have to be extremely close partners with all the EHRs who we partner with. um being able to not only pull and push all of the data into the right places is like not only table stakes. If we can't do that, health systems don't want to use us. I think the second and the reality of today is clinicians spend a lot of their days in the EHR. So much of what allowed us to win in the largest health systems was pretty direct and uh very very close partnerships with some of the largest electronic health records that allowed us to pull and push data with you know APIs that weren't ready out of the box and clinicians want to save clicks. Anytime we introduce a new product that you know adds two clicks for them in their day they're like we're not going to use it. um you know they have 15inute backtoback appointments with their patients. They're spending you know hours during pajama time doing documentation like every second and every minute counts. And so we really think about being deeply integrated into the EHR as also table stakes to actually getting real usage and adoption. And I think anything that we build or introduce, we really talk about earn the right internally a lot, which is we have to provide so much value or save so much time that people will use us. Um, but I think those are the two things that are are are close to us is we know that the product won't be used unless it is deeply interoperable and and strategically to your to your point, it's like what does EHR want to own versus us? you know, EHRs are really focused on the clinical workflows and so forth, but some of the things that we're talking about here um at least traditionally are outside of the domain where it's like, oh, connecting payers and providers together with like uh provider policies or the clinical trial matching as Janie brought up. And so these are like ent entirely we position ourselves as building this entirely new intelligence clinical intelligence layer across again providers, pharma and um payers. And so that's a it's a whole different ball game that we try to play. Yeah. It's like a different layer of scope. Yeah. I'm curious. You obviously uh are both relatively newcomers to healthcare. Um people have these like, you know, uh there's lots of futuristic healthcare AI takes of, oh, everything will look different. You know, now that you've been in healthcare for a bit, you obviously live at the edge of AI. What have you like changed your mind on around this, you know, as you think about what healthcare looks like in 10, 20 years? I mean, any any updates to your mental model from from the time actually being close to the problems? One thing that I was hesitant about before and actually it's a common thing when I'm trying to recruit engineers that people ask me around is definitely like oh healthc care heavily regulated space and and it is rightfully so you want to keep uh the patients at the end of the day safe. Um but one of the interesting things that I actually think is a that surprised me how much is coming into the companies actually there's a lot of really favorable regulatory tailwinds as well where you think about like government actually really wants interoperability between all these systems that we talked about and so agents can access this information. Um the government actually just in January the FDA released updated guidance on clinical decision support what I work on in such a way that they used to have guidance from like 2022 that required you to have like mention all these options and do all these other things but actually it's a very forward and forward-looking way. And so I think for me what's been really cool to work on is I think this there's this very special moment both in AI in general we all know that but there's a special moment also regulatory and healthcare as well. I think one thing I would call out is I think for the very reasons things are higher stakes or you know potentially considered more difficult in healthcare. I think it's where some of the hardest AI problems will get solved first just because the bar is so high. I think when I first joined I was like oh this is where we'll be on the tail end of where like all of the AI innovation will actually be able to be applied. But when you think about like zero error evalves or multi-step workflows that have really really low tolerance, I actually think a lot of the innovation will will happen here just because we have to or else we can't ship. Yeah. Cuz like in other domains, you'd much rather just solve the 80% is good enough. Yeah. 8020 doesn't work here. And building off that, I think traditionally I think there was a bit of stigma that oh healthcare companies are not that interesting from a technical perspective or I've seen that or faced that myself. But these are actually really really hard and fun problems from a pure technical perspective beyond just the impact like how do you bring the latency of this thing down um and make it really high quality. Yeah. How do you bring the latency of things down? Yeah. Yeah. Yeah. Um so okay let's answer the latency question. Um and and maybe hopefully not too redundant with some of the things I've said earlier but some part of it is with any latency you have to like what is actually what is really your bottleneck. In a lot of workflows, it's sometimes it's the model itself. And so that's where like our data flywheel, our post- training team and so forth come in. So that can you make the models far more efficient. So that's one aspect of latency. But there's whole another aspects of latency where it's like okay on top of that if you use a constellation of different models, can you use can you first use like a it's like kind of thinking fast and slow. Can you use a cheap fast model that triages and hands it off to a larger model that where you get more intelligence and so forth? And so all these clever tricks to to make it work. Um and by the way we are totally we also realize that the parto frontier is changing and so these tricks were may not get us to where we want to be in five years but we need to if we want to build a useful product right now. Should we go to the quickfire or you want to ask more about a bridging stuff everything that's not a bridge into the uh into the quickfire. But I don't mind. I I mean I I was just I think I feel like Janie was on the topic of more longtail stuff which is not the 8020 thing and that really matters and I you know if you have any tips or cool stories or just general approaches that have worked for you that's interesting to dig into. Yeah I mean I think one of them is even just how we staff our teams looks different than a traditional software engineering team. I'd say uh we have a bunch of folks with different roles who are clinicians and so we have this role called the clinician scientist and I think I heard one of our leaders refer to them as mutants recently but they are people who've had clinical background so MDs typically who are also deeply technical somewhere um on the spectrum of like basically a full stack engineer all the way to like extremely extremely scrappy prompter, but having each of these people embedded within our teams instantly raises the bar for everything that we build because not only are they determining like is this product clinically useful, but they're deeply deeply embedded in our whole eval process. And so when we talk about LFDs, when we talk about what is our actual evaluation criteria, like you don't want Chai or me creating what those are because we don't have clinical background. But I think that is probably unique to a bridge but has been gamechanging. And when you think about where the puck is going, you have people built with clinical backgrounds who are technical and where AI tools are going, they just become more and more uh critical and like the killers of the team. And so I think that's one. And then I think the second is just the scale at which we do eval to catch that long tail up front before anything ever gets into production is something that we've pretty much like really really started to fine-tune both from a scale but when do we know we need to get several hundred versus several thousand offline responses? What helps us make that quick decision and make this less of an art and as much of a science as possible? But I think that's also been something we've had to tune over time. And you have partners who basically opted in to give you those evals. Yeah. So we um work either internally or with third party for offline evals and then we have uh customers who also agree to give us whether it's like thumbs up, thumbs down to like choose this or that. Um a lot of data to get us to what is as close to to fully confident as possible. Yeah. The the term that comes to mind is kind of like active learning on things where you know you're weak. I feel like it's is kind of a lost art, but uh is a lot of the polish that comes into doing something like this. Really? Yeah. 100%. I mean, maybe on a totally unrelated note, like Chai, I think you obviously had a very uh uh story running Glean before heading over to to to a bridge. And so, you know, I'm curious like that's it's obviously was one of the early AI app success stories. Guess reflecting back on that experience like what do you think Glean got most right, maybe most wrong? Um yeah, curious for your reflections. I think the I attribute Glean's success really to uh very very strong technical foundations um that have really stood the test of time. Um and so it started with with it started with the known problem and like finding information work is hard. The best technology at the time was to build really really high quality search. I think a lot of times enterprise search startups failed because the quality wasn't great enough. But the learning that people took away from that is oh enterprise search is not good enough. And so like quality I think really changes the game of like if something can be useful or not. It's like similarly like people may have taken it away like oh Alexa voice assistants are not that useful but when you have quality things can change the game. And so I think Glean's early foundations by bringing people who had built search at Google, the best place to have ever built search, um, and being really creative and having a very concrete problem to solve, but with the right technical backgrounds laid the foundation for all of its success for the many, many, many years to come. And I think what's interesting is always figuring out like, hey, how does a company adapt in this, as we all know and we've talked many times, in this changing landscape. And so for Glean, like how do you put this context layer to the use? um has been the thing that we we've really the last few years has been the fun from the challenge that we're like you could say like that's been the opportunity for the company as well as the challenged as well. Yeah, definitely a competitive market. It feels like one at the epicenter of the foundation models and you know uh big hyperscalers. So it'll be interesting to see how it all all plays out when you think about can you build something that helps everyone at knowledge work as well is is is a is a massive opportunity. Yeah. always my mental model is like there's a few markets that are like the foundation model companies have to win or like big enough to to go after and it's probably like consumer code and that um and so it would definitely be uh be interesting to see how it plays out. I guess one thing we often think about on the investing side is you know the pace of of progress in models changes so fast and so the building patterns adjust so fast and it's always hard to figure out like what pieces of the way people are building today the infrastructure tools they use are going to prove persistent versus okay 6 months later we're doing something completely different because you know models have have improved I'm curious the stuff you use today like how how do you think about the pieces of AI infrastructure software that feel a little bit more persistent okay so generally if you take the thesis that the models are going to be more and more agent authentic I think before we had to build a lot of scaffolding around that like in previous games like I've we've effectively we made our own DSL effectively and you can view the like um uh because the models were not capable enough so you needed to simplify things um and you can view it similar to like other agent frameworks but over time if the models become more and more agentic and can use the similar tools that we already have where it's like computer use writing code itself in sandbox I think much more around I think far more about like what are the right context layers and the tools to give agents. And then I think the other things that I think about are actually how do you really build truly event- driven real time systems and especially at a bridge again where you're doing something real time in the conversation. And so I think there's a lot of event- driven technology and by the way stuff that we've always used in the past whether it's Kafka temporal sockets and so forth how do you bring that together is I think actually also durable um or thinking about patterns in which humans collaborated with each other on Google docs how do you think about like CRDT and so forth when you have conflicts when you have multi-agent systems so all these things that we've built for I actually think the things we've built for humans are actually the things that are going to be continue to be durable just with like a thousand times more the scale of agents is running at them and so yeah so make sure that they scale of course and fast and whatnot without a doubt. Yes. Does a bridge become more agentic over time then you know what what does the the next more agentic version of that look like? Um cuz you're already pretty proactive it's with like the notifications I guess. Yeah. Uh and so I view that as like a piece of being agentic. But I also view maybe some of the things we mentioned before like oh reacting to labs or like you know doing work in the background or doing even more capabilities on behalf of the clinician who we believe has a super important role to play in terms of um patient connection and so forth. Yeah. I'm curious for both of you like what's one thing you've changed your mind on in AI in the past year? I think the one I flip-flopped on and this is much more product specific is uh probably the hotter take is that prototypes are the end all be all and that purities are dead. I think we've tried switching and kind of we continue to evolve the way product is developed and uh the products that we're building are extremely complicated and nuanced and it is very very difficult for a prototype to capture the full complexity of what can we or can't we do with this data. Um what and who is this the actual right problem to be solving for in a world where software has become so cheap? Like yes this is a cool looking prototype but should we be spending any of our precious hours here? If so why? And like how does this deepen our moat in a world of decreasing moes? Um does this require custom implementation from our customer to actually use? None of that gets captured in a prototype. And so I think uh we've we're continuously evolving the way that we develop product here. But even if not written in the same traditional ways as it was 2 years ago, I think as a team, we've gotten pretty I think high conviction that in a world of so much noise, like crisp written clarity is more important than ever. It might now live in a markdown file that more teams and systems can use as context, but that's probably one that is much more function specific to me. But you're disagreeing with the consensus that purse are dead and you are like like we should partner with AI to create great documentation. But I think first probably most important is strategically answering like why is this problem the one our company and our product should solve. What happens if the next 20 competitors build this? Like why what is our right to win and does this help us differentiate in any way or are we just adding noise? I think it's important. It's a high bar. I don't know if I could answer that because a lot of the times the answer is let's do it first. And I think when the cost of doing it first is so expensive, we just talked through the the process of getting something out to customers, I think you need to have a higher bar for like as a business, should we invest here? And I think as all of our roles evolve, like one of product or like all of our jobs become like should we do this thing? And I think that's something that is worth the time spending up front on. And then um as you think about prototypes, it's still really valuable to quickly show here are the 20 ways we could do it. Clinician like I would love your feedback like which one resonates more or as you get into deeper fidelity. You can also make the prototypes deeper fidelity and like get it as close to production ready as as possible. But um beyond that to actually get it out to customers, there's a lot of implementation details, security compliance, edge cases, um things that never get caught in a prototype that I think need to be written out somewhere. And so they look different, but I think still more important than than ever. Yeah, it's interesting. I imagine a lot of that also is is like given the the context of the stage that it bridges at. I feel like for so many early stage companies, it's just a desperate race to you throw like 30 things at the wall. like please something just like resonate with my end buyer. Um and you know you find something and and that's you know why the prototype first approach is so powerful. But for you all it's like anything you're going to do is across 200 systems there's like a whole you know implementation change management side of things and you get a few big bullets to fire at like you know uh at at what you want those systems to do. And so I think being really thoughtful about that uh makes a ton of sense and maybe the uh the the prototype first takes will all grow into to your view of the world when they're when they're a bit more scale. I think the weekend demo versus it works at the largest health systems is like a massive massive gap. I don't think it means we can't go fast. Like I think this is the fastest I've built in my career um right now. And I think the compared to Loom Yeah. I think from a like t the complexity and the scale of the products we're trying to build and the problems we're trying to solve, I'd say yes, maybe I like updated a flow or like shipped a new feature pretty quickly, but if you think about some of the products we're building, we're trying to collapse prior authorization, like things that used to take 45 days across maybe 20 different touch points into one. Um, I think I'm building faster than I ever have. And so I think the thoughtfulness actually allows us just to go fast at the right things. Um, it sounds contradictory, but I think Yeah, exactly. Yeah. I, you know, it's it's interesting. I think in the when a lot of things are changing and in the AI discourse, like I think sometimes we've lose sight of things that always sort of stood the test of time like judgment and clarity always matters. It's like as an engineer sometimes I don't want a prototype. Actually I would like to see like I want the written the clarity that comes from writing and then we build that and again for some things of course where it's a small thing like yeah just the prototype that's what like don't sweat the details so I think the interesting thing the nuance that gets lost sometimes in discussion is like sometimes we need to recalibrate our judgment for sure because the cost and gains have changed but that doesn't mean we go all the way on one spectrum or the other. Yeah. Yeah. outside of your specific tool. I always like to ask this question. Any other AI tools that you guys are enjoying? Cloud code but but like that that feels like too too basic of an answer. Is all of rich engineering very pled on cloud code? Yes, very much so. You know I no cursor. Oh, we also have cursor as well. I mean I'm just checking the boxes here. Yeah, many many of the tools available but it's like you look at uh just earlier the day you see an engineer screen you see like six different you know claws running at it. sometimes the same person I've seen them on the sofa now with remote control as well on the mobile but like uh very much so like one of the interesting things for me is like as a relatively new person to companies like cloud code actually helps me onboard much faster or any of these and like I feel like I'd learn so much I do love the memes of like you know Claude's going to do this like I'd like to see Claude you know uh the venture equivalent is like I'd like to see Claude go do a company at a billion dollars pre-revenue like um well we always like to leave the last word in these conversations to to you both. And so, uh, any place you want to point folks where they can go learn more about a bridge, the work you're doing, any of the research you guys have done, whatever, the floor is yours. A couple places if you, uh, on our bridge website, we have a lot of our white papers where we've done a lot of interesting work such as like uh, reducing a hallucination. Very well presented, by the way. I liked it. Yeah, thank you. Um, you know, our science team rigorously defined what actually is a problem. And one of the interesting things by the way at a bridge is we actually have multiple uh stats professors on staff as well. Um so in that in that specific white paper Michael who's a professor at JHU. Um and so we have multiple and from that comes like very high rigor and then also our taste for design comes from really good presentation but but setting that aside and we we're going to have many more technical topics there. Please follow our Twitter account as well um bridge HQ. And then the other thing I I'll plug a little is um we have a open house of deep diving deep into AI and healthcare coming up with Andian horrors. Amazing. Uh well, thanks so much. This is super fun. All right. Thank you.

📌 文中提及的人物和组织

公司/组织: Abridge, Glean, OpenAI

关键字: clinical-intelligence ai-workflow-automation real-time-ai data-privacy model-evaluation