AI并未瓦解教育,它只是揭露了谎言 House of El - AI 2026-05-20

荣誉准则的崩塌:AI揭示的教育困境

普林斯顿大学取消了长达133年的无监考考试传统,这一决定自7月1日起生效,所有线下考试都将有教师在场监考。虽然表面上这是对AI作弊的回应,但其深层含义远不止于此。这反映出一种建立在过时世界观上的教育模式已然失效。这一事件不仅关乎AI、作弊或普林斯顿,它揭示了一个更大问题:当今教育评估的根本目标是什么?

在1893年,普林斯顿学生曾主动请愿移除监考,他们认为个人荣誉誓言比监视更有意义,并得到教职员工的认同。此后133年,学生们签署荣誉誓言,承诺未违反荣誉准则,这被认为足以维护考试诚信。然而,这种制度的执行机制是学生互相举报。数据揭示了令人担忧的现实:普林斯顿2025届毕业生的调查显示,近三分之一(29.9%)承认在普林斯顿期间有过作弊行为。更关键的是,高达44.6%的学生表示知晓违反荣誉准则的行为却选择不举报,实际举报率仅为0.4%——500名学生中只有两人举报。这表明,依靠学生互相监督的执行机制在功能上已完全失效。

Original English Source

Last week, Princeton University voted to end an 133-year tradition of unproctored exams. Starting July 1st, every in-person exam at Princeton will have an instructor in the room watching students take it. The faculty vote was a near unanimous. Just one person voted against. And the reason cited was artificial intelligence. Now, if you've seen that headline and thought, well, students were cheating with AI, so the university cracked down. That is the version of the story that most outlets are running. And it's not wrong exactly, but it misses something much more important because this isn't really a story about AI or about cheating or about Princeton. This is a story about an entire model of education that was built for a world that no longer exists. And the fact that one of the most prestigious universities on the planet just tacitly admitted that, even if they don't frame it that way exactly, should give everyone a pause. I'm El. I have a PhD in computer science and I analyze AI developments. To understand what's actually happening beneath this hype, to understand why Princeton's decision matters, you need to understand what they actually had. Back in 1893, Princeton students, the students themselves, petitioned the university to remove proctors from exam rooms. They didn't want to be watched. They believed that a pledge of personal honor was more meaningful than surveillance. And the faculty agreed. For 133 years, Princeton students took exams without anyone looking over their shoulders. They signed a pledge. I pledge my honor that I have not violated the honor code during this examination. And that was supposed to be enough. But it wasn't just a pledge. The system had an enforcement mechanism. Students were expected to report each other. If you saw somebody cheating, you were honor bound to turn them in. The entire architecture depended on peer surveillance. students watching students trust but with a distributed enforcement layer. And here's where the data gets interesting. The Daily Princonian's 2025 senior survey, this is over 500 graduating students, found that 29.9% of respondents admitted to cheating on an assignment or exam during their time at Princeton. Nearly a third. But the number that actually tells the story is this. 44.6% of seniors said they knew about honor code violations. they chose not to report and the number who actually did report a peer was only 0.4%. That is two people out of 500. The systems entire enforcement mechanism relied on students reporting each other. Functionally, nobody did.

作弊的隐形化与高回报:AI时代的检测困境

普林斯顿并非孤例。斯坦福大学在2018至2020年间收到的720份荣誉准则违规报告中,仅有两份来自学生。一位化学研究生指出,作弊已成为大学文化的一部分。斯坦福大学也因此在2023年开始引入部分考试监考,并于2026年4月投票全面允许监考。米德尔伯里学院的审查委员会发现,荣誉准则对大多数学生而言已不再是学习和生活的有意义元素,65%的学生自称曾违反过。当2025年末一项引入监考的提案被学生投票否决时,反映出学生们在承认普遍作弊的同时,又反对被监视的矛盾心态。

AI的出现并未制造更多作弊者,它只是让原有作弊行为变得更难以察觉和更高效。2012年,17%的学生通过手机短信作弊;到2026年,18%的学生使用AI提交未经修改的作业。作弊率本身并未大幅上升,但作弊的“天花板”却被极大地提升了。AI生成的文章流畅、结构严谨,若非教授有丰富的经验,几乎无法与优秀学生的原创作品区分。全球高等教育中,AI相关的学术不端行为已占所有作弊案例的60%至64%。传统剽窃案例的下降速度,几乎与AI辅助作弊案例的上升速度同步。作弊者本身没有改变,改变的是作弊工具的性质。AI作弊的检测率仅为6%,相当于掷骰子掷出特定数字的概率,学生作弊的成功率极高。

Original English Source

Princeton built a justice system where the jurors were also the defendants. And Princeton is not exactly an outlier. at Stanford. Between 2018 and 2020, just two out of 720 honor code violation reports came from students. A chemistry graduate student told the Stanford Daily that cheating had become, and I'm paraphrasing here, part of the fabric of the university. Stanford began introducing proctoring for some exams in 2023 and as of April 2026, has voted across its full governance chain to allow proctoring broadly. At Middbury College, a review committee found that the honor code has ceased to be a meaningful element of learning and living at Middbury for most students. 65% of students self-reported having broken it. When a collegewide proposal to introduce proctoring was put to a student referendum in late 2025, it failed. Only 48.8% voted in favor. Students voted against being watched while simultaneously admitting in large numbers that they cheat. So that is not exactly a policy failure on the university part. That is a culture operating exactly as its incentives predict. They voted to keep the honor code the same way people vote to keep their gym membership aspirationally. So AI didn't break the honor code. The honor code was already a very expensive fiction. What AI did was make it impossible to keep pretending. And here's a number that complicates the narrative even further. According to research aggregated across multiple surveys, in 2012, 17% of students use their phones to text answers during assignments. In 2026, 18% use AI to submit unedited work. The proportion of students submitting entirely machine generated work without any personal engagement has not dramatically increased compared with previous forms of technological cheating. What has changed is the quality of the output. A texted answer from a friend was obviously derivative. An AI generated essay is fluent, structured, and unless the professor has read thousands of similar outputs, indistinguishable from a competent student's work. The cheating rate didn't spike, the cheating ceiling did. AI related academic misconduct now represents 60 to 64% of all cheating cases in higher education globally. Not because AI created more cheaters, but because it absorbed the cheating that used to happen through other channels. Traditional plagiarism cases are declining at almost exactly the rate AI assisted cases are rising. The cheater didn't change, the cheat did. And the detection gap is staggering. Despite a 33% increase in student discipline for AI related misconduct between 2022 and 2026, an estimated 94% of AI generated assignments still go undetected. The detection rate for AI cheating is 6%. For context, that's roughly the same probability as rolling a specific number on a dieice.

监考的回归:社会压力与教育的表面止血

面对AI带来的挑战,普林斯顿的提案明确指出,AI工具和小型个人设备改变了考试不当行为的“外部表现”,使得其他学生更难观察和举报作弊行为。例如,学生在桌下看手机可能只是在看时间,但同时也可能在与AI互动。作弊曾是可见的,并且伴随着社会成本;如今,它变得隐形,而举报作弊反而可能带来社会风险。普林斯顿的报告指出,匿名举报的数量有所增加,这背后是学生们对“人肉搜索”(doxing: 通过网络公开他人个人信息)和社交媒体羞辱的恐惧。

令人惊讶的是,推动监考回归的不仅仅是教职员工,学生们也参与其中。学生政府的调查显示,大多数本科生要么支持监考,要么对此无所谓。荣誉委员会的现任和前任主席也支持这一改变。学生们宁愿被监视,因为在这样一个警务工作既不可能又充满社会风险的体系中,让他们自己去监督同伴,这个替代方案更糟糕。根据新政策,教师将作为“见证人”在考场出现,但不干预学生,这不禁让人质疑,一个曾依赖“见证人”却无人举报的旧系统,其新角色又将如何发挥作用?前院长吉尔·多兰教授称这一投票“令人羞耻但必要”。普林斯顿的提案承认,教师监考并不能根除作弊,这更像是一种“权宜之计”(band-aid: 临时解决方案),而非根本解决之道。

Original English Source

Students are essentially gambling and the house odds are spectacularly in their favor. In the UK alone, nearly 7,000 university students were formally caught cheating with AI tools in the 2023 2024 academic year, triple the number from the year before. And those are only the ones that got caught. By the way, the proposal Princeton's faculty voted on makes this explicit. It says that AI tools and small personal devices have changed the external appearance of misconduct during an examination, making cheating much harder for other students to observe and hence to report. In other words, the system relied on cheating being visible to the people sitting next to you. Somebody pulling out a cheat sheet, whispering to a neighbor, looking at somebody else's paper. AI eliminated that visibility. A student glancing at a phone under a desk looks identical to a student checking the time. As the former chair of Princeton's honor committee put it, if the exam is on a laptop, somebody can just flip to another window. The entire enforcement architecture inverted. Cheating used to be visible and socially costly to commit. Now it is invisible and socially costly to report. Princeton's own proposal notes an increase in anonymous reports of suspended violations driven by fears of doxing and social media shaming. Students weren't just unwilling to report, they were afraid to. And here's what I find most striking. It was the students who pushed for proctoring, not just the faculty. The student government surveyed undergraduates and found a majority either favored proctoring or were indifferent. Current and former chairs of the honor committee endorsed the change. The people being washed asked to be watched because the alternative being responsible for policing their own peers in a system where policing was both impossible and socially dangerous was worse. Under the new policy, instructors will be present in exam rooms as a witness to what happens, but are instructed not to interfere with students. Witness is doing a remarkable amount of work in that sentence, given that the entire previous system relied on witnesses who saw everything and reported nothing. Professor Jill Dolan, who served as dean of the college from 2015 to 2024, described the vote as a shame but necessary. And the proposal itself acknowledges that having an instructor supervising examinations will not eradicate cheating. So even Princeton is saying this is not a solution, it's a band-aid.

过时的评估模式:记忆力测试的终结

核心问题在于:考试究竟在测试什么?传统的考试模式——在没有资源的情况下凭记忆作答——是为知识稀缺且难以获取的时代设计的。那时,能否记住并复述信息是衡量学习的合理指标。然而,那个时代已不复存在。如今,信息普遍可得、即时可查。重要的技能不再是储存和复述信息,而是“知晓何物存在,知晓何时获取,以及如何评估所获取的信息”。

以我个人经历为例,我曾学习了四个学期的概率论与统计学。十年后,我无法凭记忆复述任何一个证明。但关键在于,我知晓这些概念的存在及其功能,懂得何时在解决问题时可能用到它们。我拥有知识的“地图”,即使忘记了具体的“街道名称”。这种“地图式知识”(cognitive map: 认知地图,指对某个领域知识结构、工具及相互关系的整体理解)的形成,是否必须经过艰苦的推导过程,我仍不确定。然而,这正是教育现在需要思考的核心问题。

Original English Source

And that's the part I want to sit with because I think the conversation should be much bigger than proctoring. Here's the question nobody in this debate seems to be asking. What is the exam actually testing for? The traditional exam, you sit in a room, no resources, write down what you know from memory, was designed for a world where knowledge was scarce and hard to access. If you wanted to know about Keynesian economics or organic chemistry or the treaty of Wisfalia, you had to have read specific books, attended specific lectures, and retain that information in your head. Testing recall made sense because recall was the skill. Your ability to reproduce information from memory was a reasonable proxy for whether you'd actually learned. That world doesn't exist anymore. Not because students are lazier or because institutions have failed, but because the fundamental relationship between humans and knowledge has changed. The information is now universally accessible instantly to everyone all the time. The skill that matters is no longer can you store and reproduce information. It is instead do you know what exists? Do you know when to reach for it? And can you evaluate what you get back? I'll give you a personal example. During my undergraduate degree, I took four semesters of probability and statistics. Four, that is two full years of mathematics that I would not wish on anyone I care about. I was personally victimized by Chibbishev Kulmagorov and the entire Russian school of probability theory. Chibbishev, for those who don't know, was a 19th century Russian mathematician who, due to a childhood condition that gave him a limp, couldn't pursue the military career his family intended, spent a great deal of time at his desk and produced theorems of such elegant complexity that students would be suffering through them two centuries later. I acknowledge his extraordinary contribution to mathematics. I also wish he'd gone outside more. He's dead now, so that's kind of nice. And I feel bad for even saying that, but still. But here's the thing that matters about that experience. 10 years later, I could not reproduce a single one of those proofs from memory. Not a single one. If you sat me down right now and asked me to prove trib's inequality, I would definitely fail. But, and this is the critical distinction, I know that it exists. I know what it does. I know when a problem I'm looking at might benefit from it. And because I know that it exists, I know exactly what to look for. I have the map, even if I've forgotten the individual street names. Now, is that knowledge a product of having done the hard work of having sat through those four miserable semesters and worked through the proofs by hand? Honestly, I don't know. Maybe maybe that suffering is exactly what build the cognitive map and there is no shortcut. Or maybe somebody could have taught me the landscape of probability theory, shown me what tools exist, when they apply, how they relate to each other without requiring me to derive every proof from first principle, and I'd have the same map today. I generally don't know the answer to that, but I think it's the question that education should be asking right now and mostly isn't because here's the broader reality.

知识爆炸时代:教育的重心转移

人类知识总量在不断膨胀。今天的物理学本科知识远超50年前,医学、计算机科学、生物学等领域皆是如此。成为一名合格的专业人士所需的时间不断增加,而人类寿命并未延长。在某个临界点,学位所需时间可能超过其所准备职业的实际工作年限。

面对这种趋势,教育需要优化目标,不再侧重于记忆。取代记忆的,是对领域“架构”(architecture of a field: 领域知识结构)的深入理解,以及在需要时能够有效获取具体信息的能力。这类似于使用计算器:你理解加减乘除的原理,但用计算器完成复杂计算,因为核心技能不是算术本身,而是“知道何时以及如何运行正确的计算”。将这一原则推广至教育,未来的考试将是“开卷考”,因为“书本”(信息)始终是开放的。评估的将不再是记忆力,而是“判断力”(judgment: 判断能力,选择正确工具和策略的关键)。

Original English Source

Human knowledge is expanding constantly. The body of knowledge required for an undergraduate degree in physics today is substantially larger than it was 50 years ago because there is simply more physics. The same is true of medicine. Becoming a reasonably competent doctor already takes the better part of a decade. The same is true of computer science, of biology, of virtually every field. And this only moves in one direction. Becoming a physicist a 100 years ago was not simple, but it was simpler because there was less physics to know. An undergraduate physics curriculum a 100 years from now will contain a century's worth of additional discoveries. The time it takes to become competent keeps growing, but the human lifespan doesn't. There is too much cannon. At some point, the degree becomes longer than the career it's supposed to prepare you for. I want to be clear here that I'm offering this more as a thought experiment than a prescription. I don't have it fully crystallized. And I think the honest thing is to invite you to think about it alongside me rather than pretend that I've solved it all. But I do generally think that as the sum of human knowledge continues to expand, the probability that any single person can meaningfully engage with the entirety of even their own field gets lower and lower. And that changes what education should be optimizing for. The question is what replaces memorization. And I think what replaces it is something closer to the way I actually use my education. Now, knowing the architecture of a field well enough to navigate it, combined with tools that let you access the specifics when you need them. Think of it like a calculator. You know how to add, subtract, multiply, and divide. You have the foundational understanding but you use a calculator for the heavy computation because the skill isn't arithmetic. It's knowing what calculation to run and when. Now extend that principle to education more broadly. An open book exam except all exams become open book because the book is always open. The skill being tested isn't memorization. It's judgment.

AI时代的人才评估:从记忆到认知架构

这种新的评估方式并非空谈,而是我在招聘数据科学家时实际采用的方法。我提供复杂的场景、真实模型输出、混乱数据和模糊需求,并允许应聘者使用任何工具。我测试的不是他们能否凭记忆写代码,因为在2026年,大语言模型(Large Language Model: 基于海量文本训练的 AI 系统)几乎可以编写任何代码。我测试的是他们能否理解如何处理AI的输出,能否在看到模型结果后提出下一个正确问题,以及他们是否具备有效的“认知架构”(cognitive architecture: 指个体处理信息、思考和解决问题的方式)来指导工具使用。这需要一套与记忆截然不同的技能,也更难以伪造。任何人都可以提示LLM编写逻辑回归代码,但并非每个人都能判断LLM的输出是否荒谬。这正是我所招聘的能力。

2026年5月,高德纳咨询公司(Gartner: 全球知名的信息技术研究和顾问公司)对350家年收入至少10亿美元的全球企业高管进行调查,发现约80%部署AI或自主技术的公司都裁减了员工。然而,裁员与投资回报率之间没有任何关联。那些裁员和未裁员的公司,其业绩基本相同。真正与高投资回报率相关的是“人员增效”(people amplification: 利用技术增强员工生产力而非取代其工作),即公司投入资源提升员工技能、重塑角色和运营模式,让员工驾驭和扩展技术。那些提升了回报的公司,并非消除了对人的需求,而是放大了人的价值。

Original English Source

This isn't just a theory about education. This is something I actually practice. When I hire data scientists for my team, for example, the assessment looks nothing like a traditional exam. I give candidates a complex scenario, real model outputs, messy data, ambiguous requirements, and I tell them to use whatever tools they want, any of them. I am not testing whether they can write code from memory because in 2026, a large language model can write almost any code. What I am testing is whether they understand what to do with the output. Whether they can look at a model's results and know what question to ask next. Whether they have the cognitive architecture, the map to direct the tools effectively. That requires a fundamentally different skill set than memorization. And it is substantially harder to fake. Anyone can prompt an LLM to write a logistic regression. Not everyone can look at the output and know that it's nonsense. That is what I'm hiring for. If that's actually something you'd be interested in as a standalone video, how AI is changing the way technical hiring actually works, let me know in the comments below. And the data supports this approach as well. A May 2026 Gartner survey of 350 global executives at companies with at least $1 billion in annual revenue found that roughly 80% of organizations deploying AI or autonomous technologies had reduced their workforce. But, and this is the finding that should be on every university president's desk, those workforce reductions had no correlation with return of investment. The companies that cut people and the companies that didn't were seeing essentially the same results. What did correlate with higher ROI was what Gartner calls people amplification. Companies that used AI to make their employees more productive, investing in skills, roles, and operating models that let humans guide and scale the technology. The organizations that improved returns were not those that eliminated the need for people. They were the ones that amplified them.

教育的本质:培养有用的专业人士

普林斯顿对AI的回应,是在考场增加一名“人类观察员”,这相当于在一个存在结构性问题的建筑中增加一名保安。他们问的是“如何阻止学生在考试时使用AI”,而不是“如何重新设计评估方式,使AI成为学生学习使用的工具”。这两个问题中,一个能让毕业生为即将到来的世界做好准备,另一个则固守着19世纪设计的考试形式。

我理解学生作弊的压力:截止日期、工作量、激烈的竞争。但教育的本质是为了你自身的成长,培养的是一套技能:思考、推理、知道自己不懂什么以及如何找到答案的能力。最终,每个学生都必须做出选择:是只为了文凭而获得工作,还是想在工作中真正有用?是追求进步,还是希望没人发现你的不足?文凭或许能让你敲开大门,但如果只有文凭而无真才实学,你将寸步难行。在一个AI能完成基础工作的世界,若要保持竞争力,你需要带来AI无法复制的东西:良好的判断力、情境理解能力、知道提出什么问题的能力。这需要真正学习过。

一个没有文凭却掌握了这些技能的人,在许多实际层面可能比拥有文凭却一无所获的人更具就业优势。文凭让你进门,但不能让你留在房间里。普林斯顿这样的机构应该认识到,其品牌价值,即普林斯顿学位对雇主的意义,并非那张纸,而是持有它的人的素质。如果评估模式如此破碎,以至于很大一部分毕业生未能真正获得学位所代表的技能,那么品牌价值就会贬值,无论是否抓住作弊行为。市场最终会识别出真相。

Original English Source

I actually made an entire video unpacking that Gardner study. What companies are getting wrong about AI and jobs, why the layoffs aren't generating the returns anyone expected, and what the companies that are seeing returns are doing differently. If you haven't seen it, I'd watch that one after this one. But just keep that info in mind and think about what Princeton just did. The university's response to AI was to add a human observer to the room. The equivalent of adding a security guard to a building with a structural problem. Instead of asking, "How do we redesign assessment so that AI becomes a tool students learn to use?" Well, they ask, "How do we stop students from using AI during the test?" One of these questions prepares graduates for the world they're about to enter. The other preserves the format of an exam design in the 19th century. And I want to be very clear. I understand why students sometimes cheat. The pressure is very real. Deadlines, workload, the competition for grades that open doors to graduate schools and employment. I'm not dismissing that. But I do think there is a fundamental truth that gets lost in the mechanics of it all. Education is something you get for you. You are usually paying for it. At Princeton, you're paying quite a lot for it as well. And the purpose of that payment is not a document at the end. It is a skill set. It's the capacity to think, to reason, to know what you don't know and figure out where to find it. And ultimately, this is a decision every student has to make for themselves. What kind of a professional do you want to be? Do you want to get the job because you have the credential, or do you want to be useful in that job? Do you want to progress or do you want to spend your career hoping that nobody notices? Because credentials might get you through the door, but if all you have is the credential and nothing much behind it, you're not going to get very far once you're inside. If you want to be competitive, and I mean generally competitive, in a world where AI can do the baseline work, you need to bring something AI cannot replicate. Good judgment, context, the ability to know what question to ask, and that requires actually having learned something.

构建值得信任的未来:AI时代的教育转型

1893年,普林斯顿学生要求不被监视,相信个人荣誉比监视更能保障诚信。2026年,他们的后继者却要求恢复监视,并非因为他们失去了荣誉,而是因为保护荣誉的系统已无法运作。解决之道并非增加更多监考人员,而是“构建值得信任的东西”(building something worth trusting: 建立一个基于信任而非监控的系统)。

我们需要一种评估模式,不再将AI视为敌人,而是将其视为毕业生未来职业生涯中将普遍存在的环境。这种模式测试的,不是你在压力下能回忆起什么,而是你能如何利用所有可用工具来解决问题——这是一种适用于“无限书本时代”(age of the infinite book: 指信息无限可查的时代)的“开卷考试”。普林斯顿用了133年才承认旧模式已失效,现在的问题是,需要多久才能建立新模式。

同时,普林斯顿并非唯一一个错误应对变化的机构。许多公司也在犯同样的错误,以AI的名义裁员,而数据显示这并没有带来预期的效果。研究表明,最昂贵的AI错误与技术本身无关,而是与公司尝试用AI复制劳动力而非增强劳动力有关。真正的价值在于“人的增强”(human augmentation: 利用技术提升人类能力),而非取代。

Original English Source

I think a person who acquires that skill set without a diploma is in many practical aspects more employable than a person with the diploma who didn't learn anything. Credentials get you through the door. They do not keep you in the room. And institutions like Princeton should recognize that their brand, the thing that makes a Princeton degree means something to an employer, is not the piece of paper. It's the quality of the person holding it. If the assessment model is so broken that a significant proportion of the graduates haven't actually acquired the skill the degree is supposed to represent, the brand degrades whether anyone catches the cheating or not. The market figures it out eventually. It always does. Back in 1893, Princeton students asked not to be watched. They believe that personal honor was a better guarantor of integrity than surveillance. In 2026, their successors asked for their surveillance back, not because they had lost their honor, but because the system designed to protect it had become impossible to operate. The answer to that is not more watchers in the room. It is building something worth trusting. Again, an assessment model that doesn't treat AI as the enemy, but recognizes it as the environment graduates will spend their entire careers operating in. A model that tests not what you can recall under pressure, but what you can do with every tool available to you, the open book exam for the age of the infinite book. Princeton took 133 years to admit the old model was broken. The question now is how long it takes to build the new one. But Princeton isn't the only institution reaching for the wrong lever. Companies are doing the same thing, firing people in the name of AI, and the data says it's not working either. I made a video about what the numbers actually show when companies try to replicate their workforce instead of augmenting it and why the most expensive mistake in AI right now has nothing to do with the technology. That's the one that I would watch next. Thank you all so much for watching. Subscribe and I'll see you on the next

📌 文中提及的人物和组织

公司/组织: Princeton University, OpenAI

产品/模型: GPT-4o, Large Language Models

关键字: ai-impact education-assessment honor-code knowledge-acquisition future-of-work