可信AI的基石:金融服务中的确定性验证与溯源架构
概率机器的信任困境与合规隐患
当前金融服务领域对人工智能(AI)的应用存在一个核心误区:过度关注“提高产量”(Token Maxing)即如何生成更多内容,而忽视了最关键的可信度(Trustworthiness)和可验性(Verifiability)。现有的评估方法(Evals)并不能消除大语言模型的随机性,因为大语言模型本质上是概率预测机器(Probability Machines:基于概率计算预测下一个词的非确定性系统),无法通过测试直接转化为确定性系统。在金融领域,AI的爆发将原本耗时的“撰写问题”转化为了更难解决的“校验问题”。
长期以来,华尔街分析师依赖彭博社(Bloomberg)和FactSet来获取数据,其深层心理是为了责任分担(Culpability Displacement:通过采用行业通用标准工具来规避因数据错误导致的主观责任)。然而,这些工具都是只读的,分析师仍需手工将数据搬运到Excel模型中,导致工作时长超负荷。由于金融行业面临证券交易委员会(SEC: Securities and Exchange Commission)和货币监理署(OCC: Office of the Comptroller of the Currency)的严密监管,任何交易决策都必须有可追溯的原始数据支持,因此,无法提供确定性结果的AI在实际业务中难以被真正采用。
Original English Source
Hey everyone, thank you for being here. My name is Venu Ganesh and I'm the CEO and co-founder of Kepler. Today, I'm going to talk to you about how we built verifiable AI for financial services. First, a little about me. My career has been working in fairly difficult places to work in terms of numerical accuracy and verifiability. Began my career at Palantir where I led the compute platform as well as a lot of our USG engagements. Built and sold a alt data startup. Then was head of business engineering at Citadel. I have some Citadel colleagues here in the audience as well. I've advised a bunch of startups and been lucky that all of them have reached pretty positive outcomes. So this whole talk is a distilled version of a case study Anthropic did on Kepler. If you scan that QR code, we're the only company they've ever done a case study of, which is kind of cool. And so this is a distilled version of that. That has a lot more information about how we were able to do what we do, why it's been pretty impactful in financial services, and really where we go from here.
So my only goal with this whole talk is to convince you that AI is going to start doing some very powerful things in terms of producing work product. So everything that I tell you is inevitable. Whether it's Kepler or whether it's anyone else, we are already on this journey and this trajectory. So we all better be ready.
So I first want to observe every talk that we've seen in the financial services space so far has been about producing more. Like how do I token max? How do I get AI to do more? Very little of it is about how to actually make AI trustworthy or produce trustworthy products. And what's interesting is even terms like trust and verify-ability have actually been largely abused. Evals are not verifiable. You cannot take a non-deterministic LLM and eval your way to something deterministic. These are probability machines. And the kind of underlying reason for this is that AI has made a writing problem a reading problem. We can produce insane amounts of content, whether that's code, whether it's marketing, whether it's like a DCF in record time. But we can't easily verify this. And that's because for years the hardest part about this whole process was actually producing the work. Edge and alpha came from people like Citadel being able to hire hundreds of analysts who could scour the internet and understand where any source of alpha could exist. So it really came from this idea of being able to consume content. The problem is when a model reads everything, you have the most real version of alpha decay that you possibly can. There is no edge if everyone can look at Tegus and get all the same information.
So the hard part now is trusting what actually got produced by the model. And this is kind of funny. This is not necessarily a finance problem. Every system that exists has some form of this. Software, we run CICD. We do unit tests. We do integration tests. We have code reviews. When a doctor writes a prescription, a pharmacist fills that prescription. So if it says 10,000 milligrams of a medication, someone catches that. We have a pilot and a co-pilot. We have an EMT that's a primary EMT and a secondary. In finance, we have maybe an overworked VP as a verification layer, but that concept doesn't really exist.
And now I'm going to say something even more aggressive. The reason that people buy products like Bloomberg and FactSet is to displace culpability. When you buy a tool like that, you know that information free. It exists in SEC filings. But, you believe that because a bunch of contractors or folks overseas vetted this data and stuck it in a central instance, at least if it's wrong, everyone on Wall Street has the same incorrect information. And that's interesting in certain ways. It's also kind of a scary proposition. And that's because every one of these tools is read-only. You as an analyst look at Bloomberg, you look at FactSet, you consume information, and you produce the work product. And that's where the biggest gap and biggest wall to real meaningful adoption across Wall Street has actually been.
And so, why does this actually matter? We can't use AI properly in this ecosystem. There is a reason that analysts are still working till 4:00 a.m., and there's a reason folks in investment banks are actually dying because of the hours they're putting in. Because AI can produce a very confident answer, but when it comes to producing any kind of meaningful work product, we're totally lacking. Not only that, we have the SEC, we have the OCC, we have a number of these regulatory agencies who exist to make sure that you are not insider trading, or you're not producing a trading decision that can't be justified or backed by some set of primitives or some set of information that you can reliably say, "I made this decision because of these sources of information." And so, in this ecosystem, the biggest challenge is how do I get AI to jump from producing the search technologies that it's doing right now to producing work product. That can mean a fairness opinion, it can mean a DCF, it can mean an investment memo, it can mean looking at every SIM that your firm had 5 years ago and figuring out why your IRR number wasn't what it should have been, or anything else.
确定性基底:溯源与验证的底层逻辑
在建立这种合规与信任的防线后,要在金融等高精度要求的场景中安全使用人工智能,必须为其配备确定性基底(Deterministic Substrate:一种通过传统代码或数据库等确定性系统构建的运行保障层)。目前行业通行的做法是“引用来源”,但简单的网页引用或事后审计(After-the-fact Audit)并不能解决数据本身的真实性。
真正的数据校验(Data Verification:利用可重复、数值可验证的机制来证明数据的确定性正确)与简单的文献引用有本质区别。例如,当从10-K文件(10-K Filing: 美国上市公司向证监会提交的年度财务报告)中提取营收数据时,系统必须能确定性地证明其准确无误。金融行业的一大特征在于:两个投资经理即使基于完全相同的数据,也会因为各自机构的本体论(Ontology: 机构内部特定业务术语与逻辑关系的体系定义)和投资哲学的差异,做出截然相反的买入(Long)或卖出(Short)决策。因此,数据验证不仅是寻找所谓的“客观真理”,更重要的是验证AI输出的过程是否严格遵循了组织内部的规则和概念定义。
Original English Source
And so, the whole industry has solved this by citing things. When you search the internet with Claude or ChatGPT, it gives you a list of sources that pulled information from. What's ironic is you can't easily curate those sources. So, if you find my random Substack that says Palantir is going to be $3,000 a share, and you trade off of that, please go do that because it'll be very helpful for me. But, that's not a real vetted source. Seeking Alpha and some of these blogs or Reddit posts, they're data points, but they're not real vetted sources. So, showing where you got the information from is only half the battle. And that's where we're limited right now. The citation is effectively an after-the-fact audit.
Now, a verification is a deterministic, repeatable, numerically verifiable mechanism that we can use to produce validity that a number is right. That was a lot of buzzwords. A verification just means I can prove deterministically this number is right. So, when I extract a revenue number from a 10-K, I can deterministically prove that is the correct number from that 10-K. So, these are two sides of kind of the same game. One is showing where you got the sources from, and one is verifying that the information that you pulled out is actually correct.
This becomes challenging. It becomes challenging because verification is not a outcome. It lives in the path or the set of steps that are required for you to produce information that is valid for your individual firm. Finance is one of the rare industries where two people can be have the same information, and one can be long a stock, and one can be short that stock with the exact same data. And so, the idea of verification is not actually ground truth. It is verifying that you got an output that respects the nouns and verbs or the rules of your organization. If a desk at Citadel, like a TMT desk at Citadel, believes that a particular stock is going to go up into the right, another TMT desk at Citadel may have the exact opposite belief, and they may have the exact same verification mechanisms that produce vastly different outcomes. And so, this really comes on to something simple. What are the sources that we trust? What are the transformations that we apply on those sources to produce information that we care about? And how do we make sure that's codified in a way that makes logical sense? And so, the whole point of this is simple. This came from a 10K is not the the validation as a whole. Your job now is to figure out as an individual, how do I use AI in a verifiable way? And here's the honest answer. You don't. AI is great at doing non-deterministic tasks. It can solve problems in a way that's novel. It can figure out exactly how to do EBITDA adjustments, but it can't be the one responsible for doing the mathematical adjustments. Because it's a probability machine. It is great at next token prediction. So, the second contention of this whole talk is that you cannot use AI to produce verifiable work product in finance without augmenting it with a deterministic substrate. Which effectively means, if you're a portfolio manager at Citadel, you have access to a number of deterministic tools that you use to make your trading decisions. We need to model AI like that PM. It needs to exist in grounding, needs to exist with a certain risk threshold and verifiability.
智能分工与推导链的工程实践
在厘清了确定性验证的边界之后,Kepler平台通过三大核心支柱,实现了高精度的AI金融分析:
- 原子级溯源(Atomic Provenance:将数据引用锁定在最细粒度的源头,禁止模型擅自修改数值):模型在工作时仅生成指向源数据的引用指针,而不直接接触、修改或运算该数值本身,具体的写入操作交由高保真的底层数据库执行。如果提取出的数值无法通过确定性校验,系统会直接将其过滤,绝不将其呈报给用户。
- 范围确定性(Scope Determinism:严格划分AI的逻辑推理范围,将计算与推理分离):由于大语言模型在逻辑规划和分析理解上极其强大,但在数学计算上极易出错,因此平台将非确定性的推理逻辑与确定性的计算彻底剥离。模型仅负责规划“需要计算什么”,而将具体的加减乘除、指标计算交给高效的CPU指令执行,从而大幅降低了计算成本。
- 推导链(Derivation Chains:记录并回溯每个衍生指标的计算过程与逻辑链条):由于不同金融机构对企业价值(Enterprise Value:公司总价值的衡量指标)或息税折旧摊销前利润(EBITDA: 息税折旧及摊销前利润)等指标的计算公式大相径庭,平台会记录下生成每一个数字的链条。这就像人类分析师的工作草稿,支持随时回放和倒带,从而确保了即使是最复杂的调整后财务数据,也具备完全的透明度与可追溯性。
Original English Source
So, here's how we did it. Here's what Anthropic was excited about. We have three tenants that we use to ensure that numerical accuracy is a tenant of the Kepler platform, and that allows us to produce pretty powerful, pretty verifiably reliable information. The first is atomic provenance, and I'll talk through all of these. The second is scope determinism, and the third is derivation chains, uh which we'll all talk about.
But the core premise is this. There are certain things humans should never do. And I will tell you right now, I don't believe a human should sit there and look at a PDF 10K, 10Q, 8K, or earnings call transcript and on one screen take that number and put it into an Excel model on the other screen. Humans were not built for that. That's not like there's a reason we don't have databases just written in our heads. And so, let's break this down for how you can actually use AI to produce work product.
Let's talk about provenance. Provenance as a whole just means writing down exactly where you got the information. Now, uh have a lot of respect for everyone else who got here on stage, but we did a model we trained a model that was really good at extracting information. It outperformed foundation models, it's 94% great. It's in the article. Who here would trade off of something that's 94% accurate? Right. So, fine-tuning your way on probabilistic solutions still is it's really cool and like TechCrunch will be really excited about it, but like no one else really cares about it. And that's because a wrong number is still wrong if you're in that unfortunate 6%. So, with atomic provenance, what we do is the model writes effectively a reference to the number. It cannot write the number or manipulate the number in any way. It doesn't even understand what that number is. We have tools that are really good at understanding numbers. They're databases. They are systems that can codify information and read them and write them with appropriate fidelity. So, that's the first piece of how we do things. Anytime a model makes a decision, it makes the decision to figure out exactly where it got the number from and hand off to something that can write that number. We then run it through a deterministic check where any kind of a wrong number, if we can't verify it independently, we strip it out. That number will never make it to someone if it doesn't follow the deterministic check, the Providence ledger, and most importantly, uh the whole cycle kind of repeating at least a couple of times. So, this is not me saying, "Have 10 models and each individual have OpenAI check chat check Anthropic and have Anthropic check XAI." These are not probabilistic systems evaluating each other's work. There's a core canonical process of extracting a number, persisting it, and making sure that process actually occurred properly. And this is when I say atomic traditional atomicity in like database land.
Second, scope determinism. This is This was a super controversial idea like a year ago, and VCs were like, "This is crazy." Now, this half of this talk has been This session has been about this. Um the model is really good at reasoning and planning. Intelligence is commoditized. GLM 52 shows it. You can download it off of Hugging Face right now. You have something as powerful as Opus 48. Now, what the model cannot do is math. And why would it? Why would I run 1 + 1 through a multi-billion parameter model instead of one CPU cycle? Unless you're companies that are giving bonuses on people token maxing, which is another Polarys thing. [snorts] Um and so, what the model does is the model decides what to compute. It never does the computation itself. And so, from the Kepler platform perspective, what we do is we split the deterministic pieces of the model, which are none, from the non-deterministic pieces of the model, which are all of them. And we give the model the right tooling and technology to calculate the deterministic pieces. Now, let's get really concrete. A model can read something like, "Okay, I need to understand what net margin is. I know the right pieces of information to go to to get that data, but I can't be the one responsible for doing the code behind the scenes to pull that number out of a PDF or parsing the XBRL behind the scenes to pull that information out. That is code." Now, the deterministic pieces of the platform pull that information out and persist it outside of anything the model understands. And with those two together, we can actually produce a numerically accurate answer to the question of like, what was the net What was the net margin of this stock or this company last quarter? This also is a lot cheaper because again, I don't need the model to do a bunch of stuff it shouldn't be doing.
Now, the last piece here is how do we actually do reconciliation? When someone asks about a ratio like a gross margin or anything else that doesn't exist in a filing, therefore, I can't just go look up what the gross margin is or what the set of EBITDA adjustments were. The other thing that's really complicated is everyone calculates these ratios and these multiples differently. Everyone does enterprise value calculations differently. Things that are considered recurring or non-recurring may be unique. So, not only do we have to codify that in the processes that are run, but we need some kind of a chain of events to figure out what went into producing an individual number and what went into producing an outcome. This is not any different than the chain of events that an analyst does on a desk at a hedge fund to make a risk-reward or a trading decision. It's just done by the model in a way that can be replayed and rewound.
金融AI的个性化与未来演进
通过将非确定性的生成式AI与上述确定性规则相结合,金融机构能够在数秒内完成财务报表合并(Financial Statement Consolidation:将多个子公司的财务报表合并为集团统一报表),且报表中的每个数字均可精准回溯至其最原始的财报出处。
这种自动化的工作流程将彻底颠覆依赖外包团队手动录入数据的传统模式,为高薪分析师节省出大量原本耗费在听财报电话会、整理8-K文件(8-K Filing: 上市公司重大事件临时报告)等重复性事务上的时间。这一架构不仅适用于金融,还能平滑地推广到Harvey或Legora等法律AI系统以避免虚假案例引用,以及通过美国国立卫生研究院(NIH: National Institutes of Health)白皮书进行药物分子结构提取等高精度行业。正如电子商务在迎来安全套接字层协议(SSL Protocol: 网络加密传输协议)之前 TAM(总可寻址市场)为零一样,当前的AI产业正处于类似的“前SSL时代”。唯有通过确定性验证机制,AI才能真正跨越“生成内容”的阶段,成为可信赖的生产力工具。
Original English Source
And so, that leads us to something fairly simple here. Uh we have a system that knows which data points it's allowed to produce, meaning from structured filings, from numerical data, from any other ecosystem, it knows what it's allowed to produce, and it knows what it's never allowed to produce. So, if it's pulling things out of prose, raw tables, or anything else, it doesn't do that extraction. What this allows us to do is this allows us to do things like consolidate financial statements in seconds with every number tied back to its individual source. Which means not picking any companies, but the CapIQ's, the Deloopa's that are all using contractors for this, we don't need to do that anymore. We can actually create a financial model in a numerically accurate way that allows you to build work product. We can build a DCF in the format that you actually want. In knowledge, the model will not hallucinate that a row exists that shouldn't exist.
Now, the really crazy thing and why I say this is inevitable is everything that I'm telling you generally generalizes past finance. We're picking numbers here because finance cares about numbers. But you can imagine a world where every court case, if you're a Harvey or a Legora, runs through the same process. A pre-processing step that understands that we can extract entities like case A versus case B and store that deterministically so we don't hallucinate citations. Or every drug formulation that exists in NIH white papers such that we never miss a compound or anything else.
And so, the kind of interesting dimension that we're entering is we're an ecosystem right now where we're almost like pre-SSL in the e-commerce ecosystem. Where like what's the TAM of e-commerce? Trillions, but how many people were comfortable putting their credit card number on the internet before there was security? Zero. So, we're in the last step. AI can now produce verifiable work product across a number of industries. Meaning the rag platforms of the past are really, really helpful and really cool in codifying workflows, but there's a reason that the foundational labs are going after every one of these. There's a reason that Claude for science is not a deterministic system, but still a rag-based system. And so, the piece that should be really exciting about this whole thing is there is a piece of this that no one has built yet in a variety of disciplines and a variety of industries. We're really good at consuming tokens. In fact, there's a club here for people that consumed a billion plus tokens. They're walking around with gold cards. It's kind of funny, actually. But, like, that's kind of hilarious, right? Like, in what time in history has an employee been rewarded for your company to pay another vendor for how much money you're spending? And so, token maxing, I think, is now being thought of as not the right approach here to actually solve your problems. So, the natural thing will become a rush to the bottom, which is an optimization problem, which we've seen over and over again. If you rewind time, I actually sat on this stage 4 years ago, maybe not this particular room, talking about like Snowflake Summit and Databricks Summit, where the idea was all of a sudden people wanted ROI on top of like to understand the ROI of their Snowflake investment or their Databricks investment. And we started doing cost optimization. We started figuring out how to make sure every dollar of capital we put into Snowflake and into Databricks went to actually producing a pipeline that people were using. That same trend is about to start. And we're figuring out right now, how do we use the right tool for the right job? And sometimes you don't need a multi-billion parameter model when one CPU cycle will just work.
So, kind of wrapping this up, AI has made producing work completely, I say nearly free, but like, thank you VCs, heavily subsidized. The reading problem is still very open. And you verifying a data point is not enough anymore. The system has to be able to track or track its own provenance and actually ensure that the numbers that you're pulling represent your own unique company philosophies. Citations got us like 50% of the way there, but the next half of the verifiability and provability is going to be how we start using this in real like valuable, verifiable work. And the interesting thing here is the work product itself is the proof. In code, we have I mean unit tests, we have every single pull request and every commit and every code review on that pull request stored in perpetuity. There are companies here trying to mine that information to create a representation of your on your ontology right now on a company specific basis. We need that same ecosystem in finance.
So, let's just say this, all these problems are solved at this point. The last remaining mile is going to be that personalization. So, the second version of this in 2027 hopefully one of us will be on stage talking about how we're now able to build verifiable ontologies that actually proxy our investment processes instead of saying how do I not spend a trillion tokens to solve this individual problem. Um so, we're also I mean Susanna, we're both from Kepler. We're growing pretty quickly. Obligatory come join us if these problems are interesting to you. And yeah, I think we're starting in finance now. We'll be in a lot of different dimensions pretty quickly. Um so, thank you so much and happy to answer any questions.
[Q&A]
Yes, your question is about provenance and from the provenance ledger, where does the number actually come from? At its core, the number comes from three different things. It comes from extracted information from the filings. It comes from either a mathematical calculation that we do to derivative production, so like a ratio or something else. Or it comes from your internal documents or other internal pieces of information that operate in conjunction with that external data to produce that individual data point. So, anytime the model effectively is responsible for telling some entity to do an IO operation, that's part of the provenance chain.
Yeah. What are your customers most excited about in your product? Yeah, so the question is um what are our customers most excited about? It's funny. Uh everyone wants AI like the dream is the AI portfolio manager. The portfolio managers don't want the AI portfolio manager. Like they want the AI analyst. And so the thing they're most excited about is a way of rapidly producing, rapidly doing kind of the repeatable painful tasks that their analysts are doing. Things like uh analyst actually sits there and listens to an earnings call transcript. If we didn't have to have an analyst do that, it would be amazing. Or an analyst sits there with like 15 tabs and they're opening every 8-K and 10-K over the last, you know, whatever years to create the V0 of a financial model. So they're most excited about getting their analyst time back. That's the honest answer. And these are expensive analysts. Like some of these folks make 6-700k a year to do this. Cool. I'm happy to answer any more questions outside as well if you want to defer it time. Yeah, the gentleman in the back.
📌 文中提及的人物和组织
公司/组织: Kepler