“你的护城河是数据模型”:盖茨基金会基于图谱与 MCP 的 Agent 检索实战 AI Engineer 2026-07-22

隐性知识:AI 时代的防护之墙

在人工智能(AI)技术飞速迭代的背景下,企业构建自身技术竞争力的核心问题在于——什么是真正具有防御性的资产?利用现代代码生成工具如 Claude Code,开发者可以极快地搭建出各种应用程序。然而,一旦将系统推向生产环境,企业就会面临监控、维护、以及下游系统依赖等一系列现实约束。更核心的挑战在于,用户的切入点到底是什么?是一个新的聊天应用,还是直接使用通用的 SaaS 产品?如果企业仅仅提供一个通用的 AI 界面,将极易被大模型厂商的直接更新所替代。

因此,企业真正的护城河并非底座模型,而是对内部业务流程的深刻理解,即隐性知识(Tacit Knowledge: 难以通过文字直接表述的非结构化经验与流程理解)。无论底座模型如何升级,将这些隐性知识建模并融入企业内部系统的过程是无法被轻易复制的。盖茨基金会近期推出的战略情报平台(Strategic Intelligence Platform: 盖茨基金会用于整合和检索企业级数据的平台,简称 SIP),正是通过将这些隐性知识进行系统化建模,为全机构约 4000 名员工提供检索支持,从而建立起持久的差异化壁垒。

Original English Source

Yes. So my talk today is about the title your data models remote. We have a enterprisewide platform that we had just rolled out here this past month. And so I'll go into details on this. I'll give you some hopefully some practical lessons here and why we made decisions we made for this uh how how you could picture your processes within a similar type framework. So first just a quick introduction. So this gets into the the title here the talk and the the framing of you know what I hope you take from this but with AI moving very fast at the frontier what what's defensible you know you can move you can build things very quickly with clogged code um but once you push things to production there's constraints you find how much of your your uh deployed stack do you want to actually own you know there's monitoring there's upkeep there's uh you know people building dependencies off your stack that you have to be prepared to handle. How much uh appetite do you have for decentralized access? This gets into I'll show you what we built, but you know what's what's the access point for users? You know, is it another chat app? Is it clawed? Is it chat GPT? Uh is it something else? What's your product differentiation from from those different SAS products? And so our team then you know with this context in mind you know thought through here you know what's our skill set here what's our competitive advantage in this environment and this is what I you really hope that you you take from this talk and you picture yourself in this but our moat here was our understanding of our internal processes the tacet knowledge that you need to to run successful AI and this is true I think no matter how good AI gets how good models get and new releases that different companies put out when when Mythos comes out or when there's a new app from Claude. Yeah, I'm not I'm not worried because the part that we've built is the defensible part that that's that's that's durable. So, these are the I I'll tell you what this means here in more detail, but these are the the processes tacet knowledge that we've modeled into what we call the strategic intelligence platform or SIP. and it rolled out here this past month in production for enterprise use across uh the Gates Foundation. So about 4,000 people.

数据湖仓:打破组织的信息孤岛

盖茨基金会(Gates Foundation)在过去 25 年中开展了极其广泛且雄心勃勃的工作,涵盖减少儿童死亡率、提升营养、发展农业以及教育等多个领域。为了衡量这些投资的社会效益,基金会需要进行高度依赖数据的洞察分析。然而,在漫长的历史中,基金会内部积累了海量的异构数据:每年有超过 2000 个赠款项目(许多项目金额超过 500 万美元),业务覆盖 100 多个国家,涉及 4000 多名员工,年资金拨付规模超 70 亿美元。这些庞大的数据长期散落在不同的业务系统和业务部门中,形成了一个个数据孤岛。

为了使 AI 能够在大规模场景下提取有价值的洞察,团队的第一步工作就是构建一个统一的数据湖仓(Data Lakehouse: 结合了数据湖的灵活性与数据仓库的数据管理能力的新型数据架构),将所有结构化与非结构化数据整合在统一的物理平台下。在这个统一的数据基础之上,团队搭建了数据清洗管道(Data Curation Pipeline),对非结构化文档进行语义切片、清洗、去重和结构化字段提取,并实施严格的数据治理(Data Governance: 确保数据资产安全、合规和高质量运行的管理机制),包括敏感个人信息(PII)的脱敏遮蔽以及用户数据访问权限的管控。

Original English Source

So first I know this is an engineering talk but the the the scope of this talk gets into data modeling internal operations processes and so I want to give very quick background here over what the Gates Foundation does because then this is what we're we're modeling. So, as you're probably familiar, the the Gates Foundation has a has a very wide scope and it's a very ambitious work that we've been doing for the past 25 plus years. And there's all kinds of, you know, broad initiatives that we're doing, whether it's for uh child mortality, whether it's for nutrition, agriculture, uh education. And these are kind of broadly the the different buckets that these different initiatives fit fit into. creating market incentives, spurring an innovation, collaboration between public and private sectors. And then the fourth one here kind of gets into the the lens that we're building here. You know, high quality data trying to derive datadriven insights from the from the actual investments, the the grants that we've put out. And over 25 years, there's a ton of structure. There's a ton of data that's developed. And trying to extract those insights at scale is difficult. And that's what we're trying to solve. So this slide here is a snapshot of the of some of the different uh of the work that went out in 2023 within the foundation. This gives you an idea. I just put this here to to show some of the structured the structure that we have that we're working across. So you have over 2,000 grants in one year. Many of these are 5 million plus uh many 100 plus countries that uh that that are targeted with these grants uh alumni. So there's 4,000 different employees of the foundation. Um you know many different strategies within the foundation the US within the US across almost all the states grantees the total annual dispersement over 7 billion dollars. And so this gives you some idea of structure that we're we're working with. And this one just finally here when I show the data model this will make more sense. But we have different divisions that that funding goes out through different divisions. And so this breaks down some of those divisions. So you can see different priorities and it'll make more sense in a second here. But global development, global health, uh gender equality, USP are just a sample of the different divisions. Okay. So the the fun stuff here now I hope the uh strategic intelligence platform so in a in a in a nutshell here structuring operational data for agentic retrieval. So we're building a knowledge graph with the idea of the agent consumer and here is an end-to-end look of what this looks like. So we we have different systems of record structured unstructured these have been siloed traditionally the so part of our team here the work has been to create what's essentially a data lakehouse putting everything under one roof. This is our internal enterprisewide data. It's also different different programmatic data that are uh outputs of different investments. Once it's there it's easy for us to consume. So we have a data curation layer that does different processing to it and then finally SIP here at the end with aentic chat agentic workflow as the UX you how users are consuming our platform and so it's a cross system semantic graph layer that agents can res across okay uh so some of this I'll try to speed through here just for the sake of time but this one is critical when you're dealing with systems of record with lots of complexity engagement is critical. This is something that we've we've found here repeatedly. We have to engage data owners to understand, you know, this tacet knowledge we're trying to to model. What's the full meaning of different fields, the structure of the data set, how do we join things together? How do we uh understand limitations, systematics of the data, safeguards, security trimmings, uh reporting conventions? You know, it's not enough just to answer it a question a certain way. You have to answer it the way that it's been answered in the past. And so, this is the comes back to the moat here. This is the procedural understanding tacet knowledge that AI needs and it's yeah it's the part that we that we own that's you know that's ours that um and that's what we're modeling here. Okay. So going back here just very quickly for this one this is the a snapshot here of different data curation considerations that we're that go into this pipeline. So you have [sighs] for different data sets whether it's structured unstructured there's different pre-processing filtering dduplication there's an order to different documents there can be uh inconsistencies across documents those need to be uh handled up front there's extraction so structured field extraction semantic chunking for unstructured documents if you have figures you need to convert this into text in some way so you can do retrieval across this uh various forms of tagging that these can form connections in your graph structure extra metadata that you create during this pipeline. Then that becomes different properties in your graph. And then the third bucket here, governance. This is a important one that I think AI makes more acute things that were that were accessible previously. They're much more accessible now with with AI. And so you have to consider this. Your risk sphere is is larger. So things like PII need to be masked. you need to reconsider different uh sensitive data classifying this um making sure that there's the right entitlements for each user who's accessing your system.

语义建模:基于图谱的默会知识呈现

在完成数据清洗与治理后,SIP 的核心技术在于利用 Neo4j 构建跨系统的知识图谱(Knowledge Graph: 用节点和关系表达实体间关联的语义网络)。图谱具有极高的灵活性,能够直观地映射复杂的物理业务模型,并将原本孤立在人事、赠款、审批等不同系统中的实体有机地编织在一起。具体而言,系统内建了三种关键的层级结构(Hierarchies):

  • 资金归口层级: 基金会内部拥有 80 多个不同的战略团队。团队的年度预算审批是通过一个有向无环图(Directed Acyclic Graph, 简称 DAG)进行建模的。在这个五层结构的 DAG 中,资金流向从最顶层的战略团队向下分解,流向不同的资产组合,最终落脚到具体的投资项目。
  • 投资管理层级: 在该层级中,每个层级都独立发挥作用。为了优化 Agent 的检索效率,团队在关系创建后,自动计算并生成了“rollup manages”(汇总管理)等衍生边,以清晰区分直接管理与通过子团队进行的间接管理。
  • 人员与组织架构层级: 映射人事系统中的组织架构、汇报线,以及人员在不同会议、项目和文档中扮演的角色(如负责人、与会者、审批人)。

通过将非结构化的会议文档进行语义切片(Chunking),并在 Neo4j 中与上述结构化实体建立物理关联,SIP 创造出了一个极其灵活的全局语义层,使得 AI 智能体可以动态地跨越不同的业务系统进行路径跳转和复杂推理。

Original English Source

Okay, so that's the overview here. The this the the data model itself. Now, this is the part I'll walk through here. There's a a nice animation here, but hopefully the takeaway is you can picture your own your own organization story within what I show here. I'll get somewhat technical but it's only to hope hope hopefully to give you an idea of how we how we solved our problem and then you can hopefully uh model this to yours as well. Graph is very flexible um practical representation of a physical model. Okay. So I'll zoom through a few of these here but the just the the entry point here we have over 80 different strategy teams. These teams have annual reviews that happen. This is how the budgeting for each year is derived. And then so we model this here in the graph. The the the meetings are where unstructured documents uh enter into this system from but then there's they have a structured connection to your other systems of record. Um what I show here is a conceptual data model. So it's flat. So you're not seeing the instantiation. The actual graph there's you know many different nodes. Cardonality is it one one to n. So the actual graph it's you know even more complicated. But for the data model itself, let's let me let me show you the first different hier. So we have multiple hierarchies that exist within what we've modeled the there's different types of hierarchies you can have. In this case, this is a hopefully you can see all this very well, but it's a it's an it's a um additive DAG. So there's a all five levels here of this hierarchy from the top to the bottom matter. So you have to consider everything together. And so then there's different rollup patterns you can do to work across this this sort of pattern. In our case, we have a in path shortcut here that connects the funding path. Uh funds to bow is where we have the um the budget for each of these different funding teams that that's stored. So we have funding what's the the internal funding teams have portfolios. These portfolios then go towards different investments. Multiple funding teams fund an individual investment. So it's a endtoend relationship there. The investments are the thing that are our product. It's our it's our our business. But internally we have funds that then prioritize different different types of investments. That's what's shown here. And so you can take this down to the transaction level or you can have different uh annualbased aggregations that you map here as well. And then from investment, there's a lot of interesting things you can do. You can map to all the different organizations and you can have different types of organizations and there's actually a lot here that is still kind of green space that we want to fill in. We have all these different observables that people have produced in the investments that we want to model here. So publish reports, products, you know, all this stuff is structured and connects to the entire uh organizational picture. So I mentioned that there's different hierarchies. This is the second type of hierarchy. At this hierarchy, each level matters in and of itself. And so it's not a a DAG necessarily. And so you can actually do things like precomputing the the some of these these different shortcuts. So the hierarchy it goes from the top to the bottom contains connects. It this is showing the investment management side of the of the organization. And there's concepts of direct team management. So one team at like team level two manages the investment. But then there's also a concept of indirect management. So the uh children below team level two still should be attributed to the team level two. And so there's different things you can different games you can play with these sort of rollups to precomputee. I don't know if you can see this, but rollup manages m is a is a a derived edge that we that we create after we create the contains and manage manages edge. So I've shown two different two different lenses for one investment. There's the funding lens, the management lens and you can model both of these here. Then within within the graph, a third hierarchy here is people. You have organizations, you have org charts, and you have people who are owners, you have people who are attendees of meetings, you have people who are uh directors. There's all kinds of different roles they have. You can model these here. You can have their their uh you know, who they report to, what their uh team structure is. And these all are structured data that connects across systems. Traditionally, they existed in just a HR source system, but they're relevant for the context of the the full story. And then that leads to this connectedness. So we have different source systems that were siloed. We to understand the entire picture for the agent to understand correctly across the structure you need to find these common B these common uh these common uh these common entities that you stitch together. And so that's what's shown here. These are different source systems but they're related quantity entities that exist there. And now the agent can traverse here and understand this pretty complicated organ organizational structure. One last part here that I haven't shown yet is the the document part. So, and this is still there's there's a lot more we can do to this part. We've just been uh ingesting one different document source so far, but this is where you combine unstructured and structured. And this gets into part of the magic here that you can model with Neo4j. But we have meetings that have documents. Documents then have different semantic sections that you can or chunks that you can uh that you can model here. You can put full text indexes across these to to aid in the different uh search and retrieval approaches for the agent. There could also just be a pure graph retrieval that that the agent does. And then all these things then connect back to your your your main organizational structure. So then as a whole this is what the data model looks like. So I've been zooming in here now. You can see the the full interconnectedness of this four different systems one graph uh one semantic layer that's exposed through an MCP then to the to the agents. And so this is the so if you think of the agent's perspective, this is the the structure that it can dynamically discover and reason across at query time. And for the developer, it's also a very cool thing because it exposes, you know, what you don't know about your your the thing you're modeling. You very soon you find out that there's a gap in your understanding or there's some data set that you're not, you know, fully including. And so this this process in in of it in and of itself is very valuable.

智能检索:基于 MCP 的接口接入与评估

当底层知识图谱构建完毕后,团队通过模型上下文协议(Model Context Protocol: 由 Anthropic 提出的用于连接 AI 智能体与外部数据源的开放标准协议,简称 MCP)将该全局语义层暴露给外部 AI 智能体。相比于传统的接口开发,MCP 允许智能体在查询时动态发现图谱结构并进行自主推理。团队对 Neo4j 官方提供的开箱即用 MCP 服务进行了深度定制与分叉(Fork),扩展了支持传递会话状态(如 Conversation ID 和 Message Number)的机制,以便系统能够追踪跨轮次对话的上下文。

为了确保智能体检索的准确性并符合基金会严苛的财务与合规审计标准,团队建立了一套系统的评估与反馈闭环

  1. 构建基准问答集: 与各业务领域的负责人合作,基于其实际业务报表标准,开发了分级的基准评估问答集。
  2. 混合检索比对: 由于图谱数据是动态更新的,评估系统在运行时直接提取当前图谱的实时状态作为黄金标准(Gold Standard),并与 Agent 运行时的回答进行比对。
  3. LLM 裁判评估: 引入大语言模型作为裁判(LLM-as-a-judge),评估单次问答通过率(Pass@1)以及答案的一致性与稳定性。
  4. 迭代优化模型: 评估中发现的错误往往源于业务概念的歧义,这会反过来指导团队优化图谱架构定义、完善模式描述(Schema Descriptions)以及微调领域规则(Domain Rules)。

目前,SIP 平台已实现极高的一致性,未来基金会正计划将其扩展至更多的企业级数据集,并探索团队层面的联邦图谱联合检索。

Original English Source

Okay, let me give you a sense here what we do with this now. So this is the I showed you the platform, the graph, but then how does this relate to AI? So we've we've connected this through MCP and I you know I discussed earlier what the what's durable, what's defensible to us. What was not defensible was the was the the chat interface was the UI and even in some cases the the general chat cases the the um you know the agent interaction and so we users themselves are included already or chat GPT and so we serve the platform where they are and so it's served here now through MCP here's an example just kind of a innocuous uh question here but Neoforj has some off-the-shelf uh MCP servers Here we've actually modified these quite a bit here. We forked it. And then there's various updates to the schema. Uh things to to pass state back to the to our system. You know, the conversation uh ids, the the message uh numbers, stuff like this we we've we've modified in these MCP tools. But so that's a general chat experience. That's one entry point. The other part that we're building right now too that's very exciting is more constrained workflow experiences. And so these can also be offered through things like co-work uh clawed chat. And you can do things like um you can have your you can have MCP apps be the the you the standard entry way that users access you know different UIs that are uh ported into your your your your chat experience and you can have different sandbox based agents that then run the the workflow. And so these are these are active things that we're working on. It helps to constrain the experience compared to chat, but it at the same time it pulls from that same knowledge graph-based uh back-end platform. Okay, I've got a couple minutes. I'll kind of speed through this, but the the way eval relate to data modeling is that as you're doing eval, you find you find gaps. You find ambiguities in your data model. you find ways in which users are asking questions that uh that are ambiguous or it's um you not it's not returning things that conform with the reporting standards. So what we've done here then is we've worked with data owners. We've we've uh built targeted uh eval questions that that they that match their reporting standards. We've separated these into different complexity tiers. One challenge is that the the structured data is constantly changing. So we have to have the graph query itself that we that we create for each of these different questions and then at runtime for the eval we we pull from the live graph and then we compare that to what the agent is delivering for that question and so that's what's shown here then there's a feedback loop that you can do for this. So as you're running an eval pipeline, an eval structure pipeline, you have an LLM as a judge, we've modeled things like pass at one uh stability. So if you ask the same question multiple times, you get the same answer back, you can use LLM as a judge to to to measure this. And then there's a feedback loop here that you you can update then your your data model. you can update your uh your domain rules, your schema descriptions to help to help uh fill those those gaps that you that you find. Then after you do this, this is a this is just some some eval reporting here that we uh we show the pass at one and the the stability for our system. So we we've gotten this very very strong. the the questions that we that we end up do missing, it tends to be things that are ambiguous in some way. And so it's not wrong, it's just that it's things that might be right but not what the user intended. So that that's kind of the constant struggle that we that we that we're working around 30 seconds here. What's ahead for SIP? So we're continue continuing to fill out our existing uh data from systems of records. So things that fit into our current data model. We want to expand the primary graph to additional enterprisewide data sets. There's a lot of there's a lot of demand for a federated graph experience. So, we have a main enterprise system, but we have specific teams that have their own data that they want to link to this. And so, we're working on how to do this uh different agentic experiences like I mentioned as well. And that's it. Uh, so yeah, please if you want to ask questions, if there's things that you want to talk about, I'll be out back or you can add me on LinkedIn here and you know, keep the conversation going. Thank you.

📌 文中提及的人物和组织

人物: Mike Phipps

公司/组织: Gates Foundation, Neo4j

产品/模型: Strategic Intelligence Platform

关键字: knowledge-graph model-context-protocol agentic-retrieval data-modeling enterprise-ai