AI代码质量现状:炒作与现实的博弈 AI Engineer 2025-12-11

引言:AI代码质量的现状:炒作与现实

非常荣幸能来到这里。我叫 Itamar Friedman,是 Kodto 的 CEO 和联合创始人。Kodto 的名字代表着“开发质量”(Quality of Development)。今天,我将分享我们以及其他公司关于 AI 代码质量现状的报告,试图探讨在技术炒作(hype)与实际(reality)之间的差距。

View/Hide Original English

I'm really excited being here. So many so much pragmatic and insight and suggestions. I was sitting there uh just just before. So I'm Edomar Freriedman the CEO and co-founder of Kodto. Codto stands for quality of development and I'm going to share uh our reports and other companies reports about state of AI code quality. Uh you know trying to uh talk about the hype versus reality which was uh like one of the uh points that were discussed here quite a lot which is awesome.

近期云服务中断与AI代码生成

在过去的三到四周里,我们不幸地经历了三次云服务中断。这些中断来自那些非常注重快速发展的公司。他们自己表示,他们正在使用 AI 来生成 10%、30%、甚至 50% 的代码,同时他们也非常关注质量。那么,这究竟是如何发生的?这之间有关联吗?我不知道。但让我们来做一些推测。

View/Hide Original English

So in the last three weeks, four weeks, we saw like three outages in the clouds, unfortunately, right? And these are coming from companies that really care about moving fast, right? They're they're they're saying themselves that they're using AI to generate code 10%, 30%, 50%, at the same time, they care about quality. So how did that happen? And is it is it related? I don't know. But let's have some I'm going to share some guess.

开发者使用AI代码的比例

顺便说一句,60% 的开发者表示,他们大约四分之一的代码是由 AI 生成的,或者在很大程度上受到 AI 的影响。而有 15% 的人表示,他们甚至超过 80% 的代码基本上是由 AI 生成或塑造的。现在,人们不仅使用 AI 来进行“代码编写”(vibe coding),实际上他们甚至用它来进行“代码检查”(vibe checking)和“代码审查”(vibe reviewing)。

View/Hide Original English

So by the way 60% of developers say that the like quarter of their code is either generated by AI or in in like uh uh shaped by I and 15% say that even more than 80 80% of their code uh is basically generated or or shaped by AI. Now people are using AI to do vibe coding but actually they're even doing it for vibe checking vibe reviewing.

AI代码审查与规则遵循情况

这是 Claude 的一个命令。这是给 Claude 的代码安全审查命令的提示。大约两个月前,这曾引起轰动。你们知道我在说什么吗?它上面写着,我不知道你们是否看到了。你是一位高级安全工程师。很好。然后,在某个地方,它写着“请排除拒绝服务攻击”。不要“捕获”拒绝服务攻击问题。也许这是导致我们出现云服务中断的部分原因。可能不只是这个,但你们懂我的意思。我们需要严格对待我们如何处理质量问题。这不仅仅是“代码编写”或类似的东西。我们有时会进行“代码编写”。

View/Hide Original English

This is the command of cloud. This is the prompt for the command of claude code for security review. It was hyped like two months ago. Do you know what I'm talking about now? It says there, I don't know if you see it. Uh you are a senior security engineer. Good. And then like somewhere there uh down the line it says please exclude denial of service. Don't don't uh catch denial of service issues. Maybe that's part of the part of the reason like we're we're having uh cloud outages. probably not just that, but you get the point. Like we need to be rigorous about how we deal with quality. It's not just like vibe quality or or so like we're doing vibe coding sometimes.

我们来看另一个例子。Cursor,我想大家可能都用过,或者 Copilot,大家可能都用过规则,对吧?我们将对此进行讨论。你在代码生成上投入了成本。一段时间后,你会明白,如果你投入了,你就能从中获得更多。我们询问了一群开发者,并且我也在问你们,请大家想一秒钟,在座的所有开发者,当你们编写 Cursor 规则或 Copilot 规则等时,你们是否觉得它们被完全遵循,还是大部分被遵循?你们知道它们被遵循的程度吗?它们在技术深度上被遵循到什么程度?我们得到的回应,从你们在屏幕上看到的来看,主要是 B、C 和 D。它们被遵循了,但并非完全遵循。

View/Hide Original English

Okay, let's go to another example. Okay, cursor I guess like or or pilot most of you use rules, right? We're going to talk about it. You invest in code generation. After a while, you understand if you invest, you'll get more out of it. And uh we we asked like a bunch of of developers and I'm asking you as well think think for a second for all the developers there in the audience like when you write cursor rules or copilot rules etc. Do you feel they're completely followed or it's like mostly followed? Do you know how much they're followed? And what extent are they followed? It's rigorously like how technical deep they're they're being followed.

这意味着我们正在生成代码,并试图将其推向标准,但它并不一定能达到我们想要的质量。我将分享更多统计数据、信息以及来自三份报告的一些见解:一份由 Codo 完成,另一份由 Sonar 完成,还有一份由 Far 完成。它们都专注于代码质量审查等。样本量涉及数千名开发者,在某些情况下甚至更多。涉及数百万个拉取请求(Pull Requests, PRs)以及数十亿行代码被检查。

View/Hide Original English

So the what we get back like the answer from what you see here on the screen is mostly like B, C, and D. They are followed but they're not completely followed. Okay. So that means like we are generating code trying to push it to the standards but it's not necessarily still like getting to the quality we wanted. I'm going to share a bit more statistics and and information and some insight from three reports. One done by Codo, another by done by Sonar, another by far. And all of them are are focused on code code quality review etc. The sample size is thousands of developers in some cases even more. Millions of pull requests and and a billion of of lines lines of code that were uh uh being checked.

例如,Sonar 是一家公司。是的,有点像在 AI 出现之前就存在的公司,但他们能够大规模地查看代码。他们会进行大量的代码检查,这些检查不一定专注于 AI,但对于从各个方向检查你的软件是必要的。这就是为什么他们的规模以及他们看到的代码规模是巨大的。

View/Hide Original English

Like for example, if you think about uh Sonar, this is a company. Yeah. A bit like coming from pre-AII, but they see code at scale and you they're doing like a lot of uh checks in code that are not necessarily AI focused, but are necessary in order to check uh your your software from all possible direction. And that's why their scaling and the scale of the code that they're seeing is is immense. Okay.

生产力“天花板”与智能代理工作流

我的目的在于分解“代码质量”的各个维度,并分享一些统计数据和见解。我想从最终结果开始。这是我希望大家在接下来的 13 分钟里能带走的关键信息。我们从代码生成开始,开箱即用,比如自动补全等。你为此投入成本,并能从中获得更多。但是,代码生成所能带来的生产力是有一个“天花板”的。然后我们转向“代理式代码生成”(agent code generation),称之为 Gen 2.0。这有一个更高的“天花板”,它能带来更多的生产力,尤其当你为此投入成本时,例如制定规则等。

随着 AI 打破 IDE 的界限,我们可以开始使用 AI 来实现“代理式质量工作流”(agentic quality workflows)。它可能在 IDE 内部,但事实是,如果你考虑组织中的所有工作流,特别是如果你有超过 100 名开发者,你很可能有很多与质量相关且需要自动化的工作流。这就是你开始打破生产力“天花板”的地方,如果你为此投入成本。最后,我认为你需要这些“代理式工作流”(agentic workflows)。持续学习。我们可能稍后会稍微触及这一点,因为质量是动态的。只有当你真正拥有动态的质量工作流、规则和标准时,你才能最终打破“天花板”。届时,你将看到承诺的 2 倍甚至 10 倍的提升,而不是像麦肯锡和斯坦福等机构所宣传的那样。

View/Hide Original English

So for example, we took information from from their report and eventually my purpose here is to break down the different dimension of what uh code quality means and give you some share some stats and and insights. I want to start with the end. Okay, this is the takeaway I want you all all like to take from from the next 13 minutes that I have. We started with code generation. We like out of the box use it autocomplete etc. and you invest in it and you can get more out of it. But there's a glass ceiling for how much productivity you can get from code generation. And then we move to the agent code generation, right? Let's call it gen 2.0. And that's a higher glass ceiling. It could do much more productivity and especially if you invest in it, for example, rules, etc. Then with AI breaking outside of the IDE, we can start using AI also for code for agentic quality workflows. It could be inside the ID, but the the truth is that if you think about all the workflows you have in your organization, especially if you're more than 100 developers or so, you probably have a lot of workflows that you are related to quality that you need to auto automate. And that's where you start like breaking through the glass ceiling of productivity. if you invest in it. And finally, I I claim that you need those agentic workflows. Keep learning. And we might touch a little bit of that like later later on, okay? Like because quality is something dynamic. So you'll only finally break break the glass ceiling if if you really have those quality workflows and rules and standard being dynamic. And then then you will see the promised 2x, let alone the 10x that you were promised the hyped. and you you heard from McKenzie and from Stanford you're not getting that I don't need to tell you that 2x 10x for that entire software development uh life cycle so

AI开发工具的市场采纳度

关于市场采纳度,一份报告显示,82% 的 AI 开发工具已被日常或每周使用。一些人(60%)报告说他们使用了超过三种工具,20% 的人表示使用了超过五种代码生成工具。如果你仔细想想,不要只考虑 Cursor、Copilot、Codex、Claude Code 等。抱歉,如果我遗漏了谁的工具,但还有 Lovable 等。它们也生成代码。而且,你将在两三年内使用多达 10 种代码生成工具,我对此有信心。稍后可以来找我谈,我会说服你。而且,这种趋势是从基层开始的。50% 的使用量来自不到 10 个团队,每个团队少于 10 名开发者。但它也在向企业传播。我敢肯定,你明白我的意思,它正在大规模地向企业传播,而不仅仅是少数几个开发者。在过去一年里,我们看到越来越多的企业开始使用代码生成工具。

View/Hide Original English

a bit about more about the market adoption uh one of the report says that 82% of adoption already for AI dev tools are being used daily or weekly uh some people at 60 60% 59 report that they're using more than three and 20% saying they're using more than five code generation tools If you think about it for a second, uh don't only take like cursor copilot, codex, cloud code, etc. Sorry if I'm insulting anyone in the that I forgot their tool, but there's also the lovable etc. They also generate code. And by the way, you're going to get to 10. I'm count on me. You're going to get to 10 tools in two three years that generate code for you. Okay, come to talk to me about later. I'll try to convince you. And and the thing is that it it's coming from bottom up. like 50% of the usage is coming from less than 10 teams that are less than 10 developers but it is propagating also to the enterprise again I'm sure you know I mean talk propagating to the enterprise at scale like not just like five developers

对AI生成代码的质量担忧

在过去一年里,我们看到越来越多的企业使用代码生成工具。平均而言,在报告中,我们发现 82% 到 92% 的用户每周到每月使用代码生成工具。在某些情况下,也许是极端的,也许不是。我们看到了 3 倍的代码编写生产力提升。但这并不意味着如果你有 3 倍的代码编写生产力,你就能保证任何质量,正如我之前所展示的。实际上,我们询问的开发者中有 67% 对所有 AI 生成的代码或受 AI 影响的代码存在严重的质量担忧。他们声称缺少处理质量的框架,如何衡量质量?这是一个大问题。什么是质量?我将在接下来的几张幻灯片中讨论这个问题。

View/Hide Original English

in the last year we're seeing like more and more enterprise using co code generation u so if like an average with within reports we saw 82 to 92% using weekly to a monthly uh code generation tools and in some cases Maybe extreme, maybe not. We saw 3x productivity boost in writing code. Okay, but that doesn't mean that if you have uh 3x productivity in writing code that you actually guarantee any quality like I presented before. So actually 67% of the developer that we ask asked have serious quality concerns about all the AI generated all the generated code uh uh code generated by AI or influenced by AI and they're claiming that they're missing the framework how to deal with quality how to measure quality. It's a big question.

代码规模增长带来的挑战

什么是质量?我将在接下来的几张幻灯片中讨论这个问题。在深入分析之前,请大家思考一下。什么是质量?我们实际上在说的是,“代码编写”(vibe coding)的危机正在转变和演变,即你完成了更多的任务,例如,某报告显示完成了 20% 更多的任务,提高了速度,并且打开了 97% 更多的 PR。最终,审查 PR 需要更多时间,审查 PR 需要多 90% 的时间。而且,有很多关于 AI 生成代码的统计数据,至少没有比每行代码的 bug 更少的说法。我并不是说 bug 更多,但即使每行代码的 bug 没有减少,你也会因为有更多的 PR、更多的代码生成等而拥有更多的 bug。这对审查者来说是个问题。所以,有人可能会惊讶,审查这些代码需要更长的时间,尤其是在“代理时代”(age of agents)——当时只需 5 分钟就可以使用 Claude Code,5 分钟后我就有了 1000 行代码。曾经,我需要几个小时才能编写 10 行 proper 的代码。

View/Hide Original English

What is quality? I'm going to talk about it in the next few slides. Okay, think about it for a second before I break break it down. What what is quality? Um so what we're actually saying that the crisis with V right coding uh viable coding we're seeing it shifting and evolving is that you're getting like more task being done like 20 some report 20% more task you know velocity and like 97 more% or so of PRs being opened and eventually it takes more time to review PR like 90% more time to review PR and by the way like there's a lot of statistics about AI generating ating code at least there's not less amount of bugs per line of code I'm not claiming that there are more but even if there's not less bugs per line of code you have much more bugs because there are much more PRs much more code being generated etc right so that that's a problem for the reviewer so it's somebody's surprise it takes more time to review these especially in the age of agents right when 5 minutes calling to cloud code I have 1,000 line of code after 5 minutes once upon a time it took me like hours to write 10 proper lines of code.

代码生成在不同场景下的适用性

现在,让我们稍微拉开视角。代码生成非常出色。它是一个“游戏规则改变者”(game changer),尤其是在处理全新项目(green field)时。你可能在几分钟前听别人谈论过它。它彻底改变了我们进行概念验证(Proof of Concept, PoC)项目的方式。但是,当你处理重负荷软件(heavyduty software)时,你就会发现,无论你是否愿意,我们都需要处理很多事情。当你服务于数百万客户时,你有金融交易;当你进行交通运输时,你必须处理代码的完整性、代码治理、审查标准、测试、可靠性等。这些是我们必须处理的问题。

View/Hide Original English

Right now, let's zoom out for a second. Code generation is magnificent. Okay? Like it it's a gamecher when you're talking about green field. You saw people talk about it a few slides a few minutes before me. Uh it it revolutionized how we do PE proof of concept uh project etc. But when you're dealing with heavyduty software then you you like it or not we are dealing with a lot of things when uh when you serve millions of clients you have financial transactions when you're doing transportation you're dealing with code integrity if you like code governance uh review standards testing relability etc.

代码质量问题的维度分析

现在,让我们将冰山水面下的部分分解为两个维度。这是其中一个维度。你可以从软件开发生命周期(Software Development Life Cycle, SDLC)的整个过程中审视质量问题:规划,然后是开发、编写代码、审查。代码审查本身是一个过程,但检查质量是代码审查过程的一部分;测试,这是质量的另一个部分;以及部署。我知道我没有涵盖整个软件开发生命周期,但这只是举个例子。其中每一个环节都会因为你越来越多地使用 AI 生成的代码而引入新的问题。

另一个维度是代码层面的问题和流程层面的问题。我不是要列出功能性问题,而是非功能性问题。你谈论的是安全性和效率,这些不一定是功能性的。我将展示一些关于这方面的数据。然后是流程层面,例如学习。如果因为 AI 生成的代码导致了一次重大中断,谁该负责?是 AI 还是拥有该代码的团队?最终,你需要学习并拥有代码。这是一个需要完成的过程。验证、移植、护栏、标准等。

View/Hide Original English

That's what we need to uh uh to deal with. Now let's break that under the surface part of the glacier into two dimensions. This is one dimension. You can look on the qual quality issues in throughout the software development life cycle like planning and then development writing code review. C code review is a bit of a process but like what you're like checking quality that's part of the process of code review testing, which is another part of of quality and and deployment. And I know I didn't cover the entire like software development life cycle but just to give you an example and each one of them like possess like introduce new problems that are coming because you're using more and more AI generated code. Um now another dimension to look at it is actually code level problems and process level problems.

AI对代码质量和项目进度的影响

当我们询问开发者,他们是否认为 AI 帮助减少了这些问题,还是实际上让问题变得更具挑战性时,42% 的人报告说,他们花费了 42% 更多的时间在解决问题、修复 bug 等方面,并且项目延误了 35%。我们谈论的是游戏,他们谈论的是延误。当然,这存在一些偏差。我们谈论的是质量问题及其影响等。但这正是他们在回答时所呈现的。当大规模使用 AI 生成代码时,我们看到一些报告称,安全事件增加了 3 倍。顺便说一句,这很合理。你还记得我们有一张幻灯片说代码编写量增加了 3 倍吗?所以安全事件也增加了 3 倍,每行代码的问题数量相同,但因为代码量增加了,所以问题也随之增加。

View/Hide Original English

Okay, I'm not I'm not opening the you know list of functional just opening the list of non-functional. you're talking about security and efficiency that are not necessarily uh functional use. I I'll show you some statistics about that. And then process level is for example learning. Hey, if you will have a a a bad outage because of AI generated code, who is responsible? Is it the AI or or the team that own that? Okay, Like you need to learn and own the code eventually. That's a process that needs to be done. verification, porting guard rails, standards, uh, etc. So, so all of those issues when they're introduced to thousands of developer that we asked them, do you think like actually AI helped to reduce with those problems or or actually made more like more challenging 42 people reported that they spend 42 more of the development time on solving issues, on fixing bugs, etc. and and they saw 35 uh% project delays. We're talking about we're talking about maybe games they're talking about like delays. Okay, there's some bias. We told them we talked about problem with quality and what's the impact etc. Um but that's what they they they present uh to when they they answer uh when when they're talking about like when you're mass using AI code AI generated code and

解决方案:AI驱动的测试与代码审查

那么,我们该怎么办?我一直在谈论问题。请帮助我解决它。让我们花几分钟时间来谈谈这个。一个显而易见的解决方案是测试。实际上,我们问了几个关于测试的问题,其中一个非常相关的说法是,人们表示当他们大量使用 AI 进行测试时,他们实际上对 AI 生成的代码的信任度翻了一番。

下一个有助于我们提高质量的方案是代码审查。关于代码审查,有趣的是,它是一个几乎能解决所有流程层面和代码层面问题的过程。例如,你可以设置你的 AI 代码审查工具,让它在你无法覆盖特定测试覆盖率水平时阻止该 PR。通过 PR,你就能处理测试流程问题。AI 代码审查实际上是你所能做的主要事情之一。使用 AI 代码审查工具的开发者表示,他们看到了两倍的质量提升,并且它帮助他们提高了 47% 的代码编写生产力。

View/Hide Original English

we see reports uh some of the reports talking about 3x more security inc incidents by the way it makes sense you remember we had a slide saying 3x more writing code so 3x more security incidents like the same amount of line of code the same amount of uh uh problems correlation so what to do with that like I talked about problems and problems and problems Okay, help help me deal with it. Like let's let's spend a few minutes on on that. So one one suspect of course is testing and actually really interesting we asked a couple of question about testing and one really relevant saying that people said that when they heavily [clears throat] use AI to on testing use AI to do testing they they actually double their trust in the AI generated code. Okay, that's one thing. The ne next suspect to help us with the quality is code review. What really interesting about code review that it's a process that helps almost with all the process level and the code level like issues. For example, you can set your AI code review tool to tell you block this PR if it doesn't cover certain level of test coverage. So through the PR you take care of the testing process problem. Okay. So code like code review with AI is actually one of one of the major things you you you can do and people that are developers that are using AI code review tool they're saying that they're saying they're seeing double the quality gain and they're saying that actually it's it helps them to uh improve improve 47% in productivity of writing code. Okay.

上下文的重要性:提升AI工具效果的关键

现在,我们来看看我们自己 AI 代码审查工具的一些统计数据。我们每月扫描一百万个 PR,我们选取了其中一百万个 PR,注意到 17% 的 PR 包含高严重性问题。顺便说一下,我们正在分析使用 AI 前后的情况。我还没有这些统计数据,但我们注意到,由于我们开始时,我们服务的大多数公司都使用 AI 生成的代码,所以我没有“之前”的数据。我们需要回溯扫描。这是一个非常大的数字。

我还要谈谈提高质量的另一件事:拥有正确的上下文(context)。这对于提供给代码生成工具、AI 代码审查工具至关重要。更好的上下文意味着更好的质量,无论你在哪里使用 AI。因此,当我们询问开发者,当你不信任 AI 生成的代码时(你记得 67% 的人对此感到担忧),他们说 80% 的时间他们不信任 LLM(Large Language Model,大型语言模型)所拥有的上下文。当我们询问开发者希望在 AI 生成的代码或 AI 代码审查工具中改进什么时,他们说第一项是“上下文”,占 33%。他们可以在许多改进项中选择。所以上下文极其重要。

我可以告诉你,作为 Kodto,我们的一个技术支柱就是围绕上下文。当我们连接我们的上下文引擎时,我们看到它是使用最多的工具,60% 的代码生成器或代码审查工具的调用是针对上下文引擎的。而且,上下文不一定只包含你的代码。它还可以包含你的标准、你的最佳实践。我们在 AI 代码审查中看到,8% 的上下文使用实际上来自与标准和最佳实践等相关的文件。

View/Hide Original English

Now a bit statistics from our own uh AI code review tool. We scan a million of PRs a month and we took one mill million of those PRs and we noticed that 17% include like high severity issues. By the way, we're now analyzing uh before and after using AI. I don't have that statistics yet, but we are noticing since we're starting uh most of the companies we serve, they use AI generated code. So that's why uh I don't have before. We need to go scan backwards. Uh and that's like a really big a big number. Another thing I want to talk to you like about uh when you're trying to improve on quality is is the foundation of having the right context that is brought to the uh code generation tool that is brought to the AI code review tool better context better quality across the board wherever you're using AI. Uh so when we asked developers when when you h when you don't trust AI generated code like you remember like 67% that like are really worried about that they said 80 80% of the time they don't trust the context that the LLM have okay and and and uh when we asked developers what would you like to be improved in your AI generated code in your AI code review tool they said the number one was context it was number one of 33% they can choose among many things to to improve. So context is extremely important.

未来展望与建议

CEO 的营销部门会因为我不吹嘘一点而生气。这是我们的上下文引擎在 GTC 主题演讲中被 Nvidia 的 Jensen 介绍的场景。他没有谈论我们的代码审查能力或测试能力,他谈论的是我们的上下文引擎,因为 Nvidia 已经认识到,AI 的质量、AI 生成的代码、审查和测试都将来自于提供正确的上下文。你需要为此投资,构建你的上下文,购买解决方案并进行投资,构建你的解决方案等。上下文需要包含代码版本控制、PR 历史、组织日志等。所有上下文都存储在那里,而不仅仅是你的代码库的最后一个分支。

我将拉开视角,开始谈谈建议和要点。接下来是什么?自动化质量网关。为此进行投资。人们在整个上午都在谈论并行代理(parallel agents)。你知道我在说什么,后台代理(background agents)。你可以利用许多这些工具和能力来构建你的质量网关。使用智能代码审查和测试。你需要一个“鲜活的”(living and breathing)文档。文档本身就是一个独立的故事。我不会深入探讨它。

我展示这张幻灯片已经三年了,我认为我会一直展示到 60 岁,它描绘了我心目中软件开发的未来。基本上,你有你的规范,然后你有你的代码。你有多个并行代理,它们帮助你改进规范、编写规范、改进代码、从规范转移到代码、编写可执行规范的测试。然后你将拥有你的上下文引擎——软件开发数据库。你将围绕质量和验证来构建你的工具,特别是 MCPs(可能是指模型上下文处理器或类似概念)。你将确保拥有稳定、安全的沙盒环境,让这些代理能够运行并执行验证和质量工作流。

所以,不要忘记,前进的道路在于质量。质量是你的竞争优势。AI 是一个工具,不是一个解决方案。不要只考虑代码生成。要关注整个 SDLC 或产品开发生命周期。我看到一些演讲者在迭代我们今天讨论的所有内容。我想告诉你,你将从中获得价值。我们在报告中看到,人们在安全性、可用性方面得到了更快的代码审查,我们刚刚在这方面得到了一个提示,因为 AI 生成的代码和测试覆盖率在一个月内可以增加三倍,具体取决于项目等。

在最后几分钟里,我想展示一下使用 Kodto 可以做什么的一小部分。你可以进入 Kodto 并定义你自己的规则。例如,几乎和你会在 Cursor 上设置的规则一样:“我不喜欢嵌套的 if 语句”。如果这是你遇到的问题,那么 Kodto 将会查看你的上下文,构建好的示例和坏的示例,然后开始构建一个专门用于捕获该问题的流程,并随着时间的推移提供统计数据,告诉你它何时被接受,何时不被接受。这样你就可以调整规则,真正了解并掌握你的标准。

View/Hide Original English

I can tell you that as Kodto one of our technology moes uh is is around context and when you connect our context engine we're seeing it as the number one tool that is being used like 60% of code generator or code review tools 60% of their calls to an MCP would be to a context MCP. Okay. And just to tell you the context doesn't necessarily need to include only your code. It could also include context to your standards, your best practices. We're seeing in our AI code review that 8% of the context usage is actually from files that are related to standards and and best practices etc. Okay, I have to CEO of Kodo like marketing will be mad on me if I don't brag a little bit. Right? So this is uh kind of like our market of our context engine being presented by Jensen and GTC keynote and he notice he didn't talk about our co code review capabilities about our testing capabilities he talked about our context engine that Nvidia checked because there's a realization that AI quality AI generated whatever review testing will come from bringing the right context to invest in that you need to to build your context buy a solution and invest in it build your solution uh etc. And the context needs to include code uh uh versioning PR history uh organization logs etc. That's where all the context sits. It's not just in the last branch of your codebase. Okay. So I'm I'm zooming out starting to talk about like recommendations and uh and like uh takeaways. So what what what's next? So automated uh quality gateways invest in that. People talked throughout the morning about parallel agents. You know what I'm talking about like background agents. You can use a lot of those like tools and capabilities to build build your quality gates. Uh use intelligent code review testing and you need a living and breathing like documentation and and what documentation means is is a story by itself. Uh I'm not going to double click on it. And and this is how I present for three years now and I think I'm gonna go all the way until age of 60 with this slide of how I think the future of software development looks like. Okay. So basically you have your specification and you have your code right and you have multiple agents parallel agents that are helping you to improve your spec write your spec improve your code transfer transfer from your spec to your to your code uh uh make tests which are executable specs right uh and and then you're going to have your context engine the software development database and you will build your tools especially MCPs around quality and verification and you'll Make sure you have environments, stable, secured sandboxes where those agents can run and and run validation and quality uh workflows. So don't don't forget like the path forward is quality is your competitive edge over your uh competition. AI is a tool. It's not it's not a solution. Okay? And don't like only think about code generation as the only thing. Look on the entire SDLC or product development life cycle. I saw one of the uh people talked um speakers and it iterate with everything we talked about today. I have uh I want to tell you that you will gain value from it. We're seeing in the reports people seeing like security availability being reduced faster code review you we just got a hit on that because of a generated code and test coverage in a month can can triple depends on on the project etc. with with the last minute I want to show like a really small piece of what you can do with codo. uh you can go into codto and define your own rule for example almost the same rule you'll put on cursor of I don't like nested ifs if this is a problem that you have but then kodto will look on your context build the good example the bad example and then start giving like building a workflow that is specifically to catch that issue and give you statistics over time when it's being accepted and when not so you can adjust that rule and really know and have visibility to to your standards.

当一个 PR 编写了几个 if else 语句,即使它是用 Cursor 或 Copilot 编写的,并且它们有一个“不要使用嵌套 if”的规则,最终当你打开一个 PR 时,你会收到 Kodto 的提示,并根据好的和坏的示例给出建议。Kodto 还会生成图表,提供 CLI 检查,检查每个规则,并最终告诉你嵌套 if 的情况,然后记录并学习你对该建议做了什么或没做什么,以便调整标准和质量。

View/Hide Original English

Okay. So when a PR is written with a few ifs and else although it was written with cursor copilot that had a rule do not do nested ifs etc. then eventually when you open a PR you will get uh codo uh uh catching that and giving a suggestion according to the good and the bad example. COD will also make a graph, give you a CLI checks like check each one of the rules and eventually tell you the nested if and then we'll record and learn what you did or did not do with that suggestion in order to adapt the standard and of the of the quality.

此外,还将有自动建议,你无需编写自己的规则。它会学习你的标准和质量,并提供给你。就是这样。我真的很兴奋能够打破“天花板”,无论是通过代码生成还是代理式代码生成。现在我们正进入将 AI 应用于整个 SDLC 的时代。最重要的一部分与质量相关。你需要为此投资。这不是开箱即用的。然后你最终会看到承诺的 2 倍提升,这可能是你向 CEO 承诺的,以便在他们给你相关工具的预算时获得。谢谢大家。

View/Hide Original English

Um there will also automated like suggestion. You don't need to write your own. It learns your your your standards and quality and offer that to you. And that's it. I'm I'm really really excited about like breaking the glass ceiling, okay, with what we did with code generation and then a jet to code generation. Now we're turning into the era of putting AI into work and through the entire SDLC. The most important part is related to quality. You would need to invest in that. It's not out of the box. Okay. And then you would see eventually the promised 2x that that probably promised to the CEO or something like that once they give you the budget for for the relevant tools. Thank you so much.

📌 文中提及的人物和组织

公司/组织: Sonar, Nvidia

产品/模型: Cursor, Copilot, Codex, Lovable