作弊是症状,而非病因
布朗大学一位教授布置了开卷期中考试,86名学生参加,班级平均分高达96分(满分100),其中40人获得满分。而该考试近二十年来的历史平均分仅在65至80之间,且本次试题难度被刻意提高。教授怀疑有异,遂宣布期末改为闭卷。消息一出,27名学生立即退课,另有9人直接消失——这更像是一场撤离而非退课潮。开卷时全班作答完美,闭卷时同一题型的平均分仅为10%。这意味着,即使把试卷交给一屋子异常聪明的金毛犬,让它们随机啃咬答题卡,其统计表现也与这些学生处于同一认知层级。Y Combinator联合创始人Paul Graham将散点图发布在X平台后迅速走红,两位Google DeepMind员工也参与讨论。在布朗任教34年的教授Roberto Serrano表示收到了数百封邮件,被舆论淹没。多数媒体报道的叙事很简单:学生用AI作弊,教授抓了个正着,散点图走红。但作弊只是症状——40名学生在六周后无法回答自己曾得满分的题目。他们把开卷考试喂给了ChatGPT,拿到了成绩,脑子里却空空如也。
Original English
A professor at Brown University gave his students a take-home midterm exam. 86 students took it. The class average was 96 out of 100, which is high and surprising until I tell you that 40 of them scored a perfect 100, which is just insane. For context, the historical average on this specific exam across nearly two decades of teaching it ranged between 65 and 80. And this version of the exam was deliberately harder than anything he had said before. Gosh, I wonder what they had at home that made such a big difference. The professor suspected something was wrong, so he told his students the final would be in person. Immediately, 27 of them dropped the course. Nine more simply didn't show up, which seems more like an evacuation than a drop off period. On the take-home, the class answered it beautifully. In person, the average score on that same question was 10%. Which only means if you handed the exam to a room full of exceptionally bright golden retrievers and they randomly guessed by chewing on the Scranton, they statistically would have performed at a similar cognitive tier. Y Combinator co-founder Paul Graham posted the scatter plot on X. It went viral. Two Google DeepMind staffers weighed in as well. The professor, Roberto Serrano, who has taught welfare economics at Brown for 34 years, said he received hundreds of emails and was overwhelmed by the response. The story most outlets are running right now is pretty straightforward. Students cheated with AI, professor caught them, scatterplot went viral, and all of that is true, but the cheating is a symptom. 40 students scored 100 on material they couldn't answer 6 weeks later. They fed a take-home exam to Chad GBT, collected their grade, and walked away with nothing in their heads.
经济信号下的理性选择
Serrano教授的课程名为“福利经济学与社会选择理论”,本质上是高级数理经济学,涉及形式化证明和严谨逻辑论证。往年选课人数仅8至30人,但今年春季有86人注册。Serrano将激增归因于一个因素:他宣布考试改为开卷。这一决定出于同情——2025年12月,一名枪手进入布朗大学教学楼,不幸造成两名学生死亡、九人受伤,许多学生对坐在教室里感到焦虑。Serrano还特意将期中试题设计得更难,认为无限时间能让学生更深入钻研。结果他得到了40个满分和96分的班级平均分。当他和助教将试题输入ChatGPT时,AI给出的答案“大致正确,但非常偏离,风格极其迂回”。一道本可用直接论证解决的问题,ChatGPT和许多学生却用了反证法,虽得出正确答案,但路径“非常牵强,显然不是人类推理的产物”。Serrano告诉班级他相信存在大规模作弊,并给予他们证明自己清白的机会:如果闭卷期末的成绩分布与期中相似,他就保留期中成绩;否则作废期中,仅以期末计分。结果平均分暴跌近50%,19人不及格,3人得零分。那些一听说闭卷就退课的学生,多数在期中得了满分。布朗大学学术准则副院长Love Wallace将学术违规描述为“几乎从未出于恶意”,常源于“瞬间决定”。Serrano不以为然:“你不能认真告诉我这是瞬间决定。瞬间决定是在杂货店拿错燕麦奶。逐行复制粘贴整个数学证明,同时无视你的聊天机器人可能在胡诌一种拉丁方言——这不是瞬间决定,这更像是一部预谋学术欺诈的长篇电影。”
Original English
Professor Serrano's course is called welfare economics and social choice theory. It's essentially advanced mathematical economics. We're talking formal proofs, rigorous logical argumentation, the kind of material that requires you to actually think. The kind of course that sounds less like a degree requirement and more like a psychological experiment designed to see how long a human brain can endure pure Greek letters before it simply surrenders. In previous years, the course attracted between 8 and 30 students, which is honestly not a lot. It's difficult. It's niche. It's not the sort of thing you take for an easy grade. But this spring 86 student enrolled. Serrano attributes the spike to one thing. He had announced the exams would we take home. He made that decision for compassionate reasons because in December 2025, a gunman entered an academic building at Brown and unfortunately killed two students, injuring nine others. Many students expressed anxiety about sitting in classrooms. The take-home accommodation was, as SA put it, appropriate. I agree with that. Well, he also designed the midterm to be harder than usual, reasoning that unlimited time would allow students to engage more deeply with challenging material. Instead, he got 40 perfect scores and a class average of 96. When he and his graders ran the exam questions through Chad GPT, the AI produced answers that were, in Serrano's words, kind of correct, but very off and with a very convoluted style. One question could most obviously be solved using a direct argument. Chad GBT and many of his students used a contradiction argument instead. It arrived at the right answer, but through a route Serrano described as very contrived and clearly not produced by human reasoning. It's essentially the classic AI signature, taking the scenic route through a burning building just to hand you a perfectly dry cup of water. So Sana told his class he believed there had been massive cheating. He gave them a chance to prove him wrong. If the grade distribution on the in-person final looked similar to the midterm, he would let the midterm stand. If it didn't, which is of course what he expected, he'd void the midterm entirely and reweigh the final. Which is the professor version of basically saying, "All right, the vibes are off. Everyone back in the cage." The results were catastrophic, and I'm sure you can guess in which direction. The average dropped by nearly 50%. 19 students failed the course, three scored zero, and the students who dropped the course the moment they heard in person, well, most of them had scored 100 on the midterm. Brown's associate dean for the academic code, Love Wallace, described academic code violations as almost never malicious and suggested they often come from a split-second decision. Serrano was not impressed. "You cannot seriously tell me this was a split-second decision," he said. Many of those students cheated all the way to 100. A split-second decision is accidentally grabbing the wrong oat milk at the grocery store. Copying and pasting an entire mathematical proof line by line while ignoring the fact that your chatbot might be hallucinating a dialect of Latin is not a split-second decision. That kind of sounds more like a fulllength feature film of premeditated academic fraud.
考试在测量什么?
这场考试本质上要求学生复现数学证明——拿一个早已被证明(往往在几个世纪前)的定理,从头重建证明过程。这相当于让20岁的年轻人进行“智力角色扮演”,装扮成18世纪的法国数学家,重新发明轮子,只为证明他们知道轮子是圆的。作为拥有计算机科学博士学位、亲身经历过四个学期概率论“折磨”的人,我对这种模式感受复杂。有人已经证明了氧气的存在——拉瓦锡在18世纪就搞清楚了,此后人类一直安心呼吸,从未要求每一代人重新推导化学。但在数学中,学生被期望复现几十甚至几百年前确立的定理证明,仿佛知识不算数,除非你亲自“受苦”过。当然,我知道证明的意义:它并非为了确认定理,而是训练一种认知肌肉——构建严谨逻辑论证的能力,没有漏洞,没有含糊其辞。这种能力会迁移到评估看似合理但实则错误的模型输出上。证明是健身房,定理是杠铃。但去健身房增肌和被强迫把一堆砖头从房间一端搬到另一端,只因监工喜欢听喘气声,这两者有本质区别。我学过图论、算法和计算复杂性多个学期,如今已无法复现大多数形式化证明,但我知道什么是最小生成树,知道何时问题是NP难的、该停止寻找精确解,知道需要特定算法时去哪里查找。我脑子里有地图,即使忘了街道名称。我知道如何在大陆上导航,只是说不出高速公路上每颗鹅卵石的确切位置。坦率地说,如果真需要知道鹅卵石的布局,我会去查,而不是哭着等毕达哥拉斯的鬼魂向我低语答案。一半的课程内容能否构建80%的地图?更短的课程加上引导式工具辅助探索能否构建70%?我真心不知道,据我所知也没人知道,因为没人进行过大规模实验。Serrano的考试测试的是学生能否用教授期望的特定方法复现证明。ChatGPT用了不同方法——反证法而非直接论证——得出了正确答案。那么,考试究竟在测量什么?学生对定理的理解,还是他们复现特定认知路径的能力?如果我问你去超市怎么走,你给了我精确坐标,但你是蒙着眼跳着弹簧单高跷到达的——你仍然到了超市。我并非说这场考试不好,它只是传统考试,测试传统技能,但在AI使这些特定技能变得极易外包的世界里,当外包成本降至零时,考试不再是对知识的测试,而变成了你是否选择使用口袋里那个工具的测试。
Original English
Now before I get into the broader implications, I want to pause on something that I haven't seen anyone discuss so far. What was this exam actually asking students to do? Fundamentally, it was asking them to reproduce mathematical proofs to take a theorem, something that has already been proved often centuries ago, and demonstrate that they could reconstruct the proof from scratch. We're essentially asking 20-year-olds to perform intellectual cosplay. We want them to dress up as 18th century French mathematicians and recreate the wheel purely to prove they know that the wheel is round. And I want to be honest with you here. As somebody who holds a PhD in computer science and who spent four semesters being personally victimized by probability theory and God knows only what other branch of mathematics, I have complicated feelings about this. Look, somebody already proved that oxygen exists. Antoine Lavoier figured it out in the 70s7s and humanity has been contentedly breathing ever since without requiring each new generation to rederive the chemistry. And yet in mathematics the expectation is that students will reproduce proofs of theorems that were established decades or centuries ago as though the knowledge doesn't count unless you've suffered through it yourself. Of course, I'm being semifiticious here. I know the argument for why proofs matter. The proof isn't really about confirming the theorem. It's about training a cognitive muscle. The ability to build a rigorous logical argument with no gaps, no handwaving. No, trust me, bro. And of course, I'm aware that discipline transfers when I'm evaluating model outputs that look plausible but are subtly wrong. I'm doing a version of the same thing. Finding where the logic breaks, identifying the hidden assumption. The proof is the gym. The theorem is the weight. But there is a difference between going to the gym to build muscle and being forced to move a pile of bricks from one side of the room to the other just because the supervisor enjoys the sound of heavy breathing. Fundamentally, there is a question of diminishing returns that I think gets lost. I studied graph theory, algorithms, and computational complexity across multiple semesters. I cannot reproduce most of those formal proofs today, but I know what a minimum spanning tree is. I know when a problem is NP hard and I should just stop trying to find an exact solution. I know where to look when I need a specific algorithm. I have the map in my head even though I've forgotten the street names. I know how to navigate the continent. I just can't tell you the exact placement of every specific pebble on the highway. And frankly, if I ever need to know the pebble's layout, I'm just going to look it up, not cry until the ghost of Pythagoras whispers the answer to me. Would half the coursework have built 80% of that map? Would a shorter program plus guided tools assisted exploration have built 70? I genuinely don't know. And to my knowledge, nobody does because nobody's running that experiment at scale. And this matters because Serrano's exam, the one that produced the scatter plot now circulating around the internet, was testing whether students could reproduce a proof using a specific method that the professor expected. Chad GPT used a different method. It arrived at the correct answer through a contradiction argument instead of a direct one. Sarana could tell it was in human reasoning, but the answer was right. So what exactly was being measured? The students understanding of the theorem or their ability to replicate a specific cognitive route to it. If I ask you for directions to the supermarket and you give me the exact coordinates but arrived there via pogo stick while blindfolded, you still got to the supermarket. I'm not saying the exam was bad, by the way. I'm just saying it was a traditional exam designed to test traditional skills in a world where AI makes those specific skills trivially easy to outsource. And when the cost of outsourcing drops to zero, which is exactly what happened, the exam stops being a test of knowledge and becomes a test of whether you choose to use the tool sitting in your pocket or not.
经济压力下的激励扭曲
布朗大学录取率仅5.5%,入学需要多年持续优异的学业表现,这些学生显然有能力。那么,为什么多数人(数据强烈表明是多数)选择把开卷考试交给ChatGPT而非自己钻研?我认为他们并非厌恶努力,而是厌恶那种感觉与自身需求脱节的特定努力。考虑经济背景:布朗大学2026-27学年总费用为97,000美元,仅学费就约75,000美元,十年内上涨近50%。在这个价位上,你买的不是教育,而是一张极其昂贵的收据。美国学生贷款债务总额达1.83万亿美元,涉及4280万借款人。欧洲模式不同,学费大多受补贴或免费,债务负担较低,但压力并未消失——欧盟青年失业率15.1%,英国16.2%,西班牙24.2%,英国超过100万年轻人处于非教育、非就业、非培训状态,为2013年以来最高。无论你被学费债务淹没,还是毕业即失业,结果都一样:文凭成了唯一重要的东西,学习成了你负担不起的奢侈品。过去30年,美国公立大学四年制学位费用翻倍,而家庭收入中位数经通胀调整后仅增长39%。费用与家庭承受能力之间的差距持续扩大,学生用债务填补。当文凭如此昂贵、就业市场如此严峻时,激励结构悄然转变:文凭成为产品,学习成为可选项。当一种工具能让你无需太多学习就获得文凭时,使用它并非不理性——你只是在回应价格信号。如果你向某人收取一栋郊区小房子的价格来办护照,就别在他们用复印机过境时感到震惊。布朗大学作弊的学生并非在做道德决策,在我看来,他们是在做经济决策,而他们所处的系统恰恰激励了这种行为。再加上这些学生正在应对真实的创伤——六个月前校园里两名学生在教室中丧生——以及为全日制学习设计的工作量,加上常需兼职的经济压力。持续学术努力的认知成本是真实的,期望所有人同时保持超人输出不能成为标准,因为这违背了“超人”本身的定义。你可以短时间冲刺,但如果系统要求你始终冲刺,你得到的不是英雄,而是倦怠和捷径。Serrano说得很好:“这项技术的问题在于,作弊成本基本降到了零。学生很容易屈服于诱惑。”他以经济学家的方式思考,他是对的。当一种行为的成本降至零,且被抓的惩罚微不足道时,该行为就会成为默认。解决之道不是增加监考人员,而是改变你要求学生做的事情。
Original English
Now, just how millennials cannot afford houses because we all have Netflix and love avocado and toast. Similarly, students are just always lazy. They're just lying face down on a pile of expensive textbooks, actively refusing enlightenment because the modern youth simply lacks the moral fiber of our ancestors who notoriously love doing advanced welfare economic proofs by candlelight. But we wouldn't be on House of L if we didn't explore the devil's advocate. So, let's give that a go for a second. Brown University has an acceptance rate of 5.5%. Getting in requires years of sustained academic performance. These students are demonstrably capable. So why did the majority of them and the data strongly suggest it was the majority choose to hand a take-home exam to Chad GPT rather than engaged with the material? I don't think they're averse to effort, but I do think they're averse to a specific kind of effort that feels disconnected from what they need. Let's consider the economic context. And I won't go in depth on this because this is not a macroeconomics video much to my shagran but it matters. Brown's total cost of attendance for the 202627 academic year is $97,000. Tuition alone is around $75,000 up nearly 50% in a decade. At that price point, you're not really buying an education. You're buying a very expensive receipt. For nearly a hundred grand a year, that receipt better be printed on 24 karat gold leaf and come with a complimentary house. Total student loan debt in the United States currently stands at $1.83 trillion, spread across 42.8 million borrowers. In Europe, the model is different. Tuition is heavily subsidized or free in most countries. So, the debt burden is lower, but the pressure doesn't disappear. It just shows up elsewhere. Youth unemployment is 15.1% in the EU right now, 16.2% in the UK and 24.2% in Spain. Over a million young people in the UK are currently not in education, employment, or training, which is the highest since 2013. Whether you're drowning in tuition debt or graduating into a job market that just doesn't want you, the result is the same. The credential feels like the only thing that matters and the learning feels like a luxury you cannot afford. Over the last 30 years, the cost of a 4-year degree at a public institution in the United States has doubled, while median family income has increased by just 39% adjusted for inflation. The gap between what education costs and what families can afford has been widening for decades, and students are filling it with debt. When the credential costs this much and the job market looks like that, the incentive structure quietly shifts. The credential becomes the product. The learning becomes optional. And when a tool arrives that lets you get the credential without much of the learning, you're not being irrational by using it. You are just responding to a price signal. If you charge somebody the price of a small suburban home for a passport, just don't be shocked when they use a photocopier to pass through the border. The students who cheated at Brown were not making a moral decision. It seems to me they were making an economic one and the system they're operating in incentivizes exactly that behavior. Add to this that these students are navigating genuine trauma, a campus that lost two students in a classroom 6 months earlier. A workload designed for full-time study alongside the economic pressures that often require part-time work. The cognitive cost of sustained academic effort is real and expecting superhuman output from all humans simultaneously cannot become the standard because it defeats the whole point of what superhuman even is. You can sprint for short periods. But if the system demands sprinting at all times, you don't get heroes, you get burnout and shortcuts. Srano framed it very well. The problem with this technology, he said, is that the cost of cheating has basically gone down to zero. It's very easy for students to succumb to the temptation. He is thinking like an economist and he's right. When you reduce the cost of a behavior to zero and the penalty for getting caught is kind of negligible, the behavior becomes the default. You don't fix that by adding proctors. You fix it by changing what you're asking students to do.
大规模研究揭示的真相
布朗大学的散点图因视觉冲击力而走红——你能直观看到作弊,即时、可读、适合发推。但这只是一门课、一位教授、59名学生。我们需要更严谨的大规模研究。2026年5月,Cheerovs、Smeirnov和Kizelchek发表了目前规模最大的本科生生成式AI使用研究,涵盖20所美国主要公立研究型大学的95,513名学生,发表于《科学》期刊。三分之二的学生曾将生成式AI用于课程作业,9%承认用于作弊。关键发现是:每日使用者作弊的可能性是月度使用者的三倍以上(26%对7%)。这本质上是一个“滑坡”式发现——你用得越多,外包得越多;外包得越多,没有它就越难完成工作。研究还发现AI采用和滥用因学科而异:计算机科学常规使用率最高(62%),其次是数学(53%)和商科(51%)。使用模式不同,滥用模式不同,测试的技能也不同,这意味着应对AI作弊的任何措施都需针对学科。一刀切的政策就像给所有疾病开同一种药。研究人员明确呼吁进行学科层面的评估改革。这篇有95,000名参与者、呼吁与Serrano轶事所揭示的相同结构性改革的论文,获得的公众关注度却远不及那张散点图。我们对故事、视觉、YouTube视频有反应。散点图戏剧化、可读、情感直接;而同行评审的《科学》论文密集、谨慎、悄然具有毁灭性——它是一页又一页干巴巴的数据,干到需要一杯尝起来像地中海海底的马提尼。散点图本质上是一份包装精美的学术八卦,一个视觉梗,一眼就能说“看这些肮脏的作弊者”。而科学论文是40页的直接蓝图,解释学校所建整座山正在滑入海洋。但蓝图无法干净地塞进一条推文,于是我们继续谈论八卦。两者都重要,都值得关注。视觉胜出的事实说明了我们作为社会如何消费证据——这本身就是一个批判性思维问题。
Original English
But before I tell you what I think should change, I want to show you what the research actually looks like when you zoom out beyond just one classroom. The scatter plot from Brown went viral because it's visually dramatic. You can see the cheating. It's immediate and legible and makes for an excellent tweet, but it's also one course, one professor, 59 students. That's not really good enough for the rigor that we need. So, let's talk about what the research at scale actually shows. In May 2026, Cheerovs, Smeirnov, and Kizelchek published what is currently the largest study of generative AI use by undergraduates. We're talking 95,513 students across 20 major US public research universities published in science, the holy grail, okay? It's the academic equivalent of getting a blessing from the Pope except with a lot more P values and fewer hats. Twothirds of students had used generative AI for coursework. 9% admitted to using it to cheat. And here's the critical finding. Daily users were more than three times as likely to cheat, 26% compared to monthly users at 7%. It's basically a slippery slope type of finding. The more you use it, the more you offload, and the more you offload, the harder it becomes to do the work without it. The study also found that AI adoption and misuse vary dramatically by discipline. Regular use is highest in computer science at 62% followed by mathematics at 53% and business at 51%. The patterns of use are different, the patterns of misuse are different and the skills being tested are different, which means that any response to AI cheating needs to be discipline specific. A blanket policy on this one makes about as much sense as prescribing the same medication for every single illness. The researchers explicitly called for assessment reform at the discipline level and I completely agree with them because it's exactly the same conclusion I reached in my previous video when we talked about Princeton. Which brings me to something that I find genuinely very interesting. This paper 95,000 participants calling for the same structural reform that Serrano's anecdote illustrates has received a fraction of the public attention that the scatterplot got. We respond to stories. We respond to visuals, to YouTube videos. A scatterplot is dramatic, legible, and emotionally immediate. A peer-reviewed study in science is dense, careful, and quietly devastating. It is pages upon pages of data so dry it demands a martini that tastes like the bottom of the Mediterranean Sea. The scatter is essentially a beautifully wrapped piece of academic gossip. It is a visual meme that says, "Look at these dirty cheaters." in a single glance. The science paper is the direct 40-page blueprint explaining that the entire mountain the school is built on is sliding into the ocean. But the blueprint doesn't fit cleanly into a tweet. So, we just keep talking about the gossip. Both matter. Both deserve attention. And the fact that the visual wins tells you something about how we consume evidence as a society, which incidentally is itself a problem of critical thinking.
认知债务与苏格拉底式AI
关于AI是否真正损害学生认知能力,文献具有提示性但尚无定论。Gerish 2025年对666名参与者的研究发现,频繁使用AI与批判性思维得分之间存在显著负相关,由认知外包(cognitive offloading)中介。MIT媒体实验室的一项研究使用脑电图监测论文写作期间的脑部活动,发现使用ChatGPT的参与者神经连接最弱,且这种模式在停止使用工具后仍持续存在,研究人员称之为“认知债务”(cognitive debt)。其他研究则讲述了完全不同的故事。Alfaria等人2026年发现,标准ChatGPT使用产生了最高的感知理解度和最低的实际学习效果——自信而无能力,巧合的是,这也是LinkedIn上每个科技网红的商业模式。然而,当AI交互被重新设计,加入他们所谓的“教学摩擦”(pedagogical friction)——苏格拉底式提问、引导性提示、迫使学生在获得答案前思考的提示——学习收益显著。另一项关于苏格拉底式AI辅导的研究发现,学生报告获得了显著更强的批判性、独立性和反思性思维支持,直接挑战了“AI必然让你思维变差”的叙事。因此,文献的结论并非“AI对学习有害”,而是“完全取决于你如何使用它”。将其用作思维的替代品,它会萎缩肌肉;将其用作推动你推理的苏格拉底式伙伴,它反而能增强肌肉。问题不在工具,而在交互设计。而目前,没有人教学生如何将这些工具用作思维伙伴而非思维替代品。布朗大学自己的生成式AI与教学委员会(其报告与Serrano丑闻同一周发布)发现,56%的本科生和67%的研究生每天或每周使用生成式AI。最引人注目的是,这些学生中的绝大多数报告担心这会影响自己的学习——他们担心失去认知能力,他们知道这是个问题,但他们仍然这样做。这就像凌晨1点盯着第三块巧克力蛋糕,大脑在尖叫“这对长期结构完整性是糟糕的选择”,而手已经伸向了叉子。
Original English
Now on the question of whether AI is actually damaging students cognitive capacity, the literature is suggestive but not yet conclusive. Gerish 2025 studied 666 participants and found a significant negative correlation between frequent AI use and critical thinking scores mediated by cognitive offloading. An MIT Media Lab study used EEG to monitor brain activity during essay writing and found that participants using Chad GPT displayed the weakest neural connectivity and those patterns persisted even after they stopped using the tool. The researchers called it cognitive debt. Other research tells a completely different story. Alfaria atal 2026 found that standard chat GPT use produced the highest perceived understanding alongside the lowest actual learning. So confidence without competence which coincidentally is also the exact business model of every tech influencer on LinkedIn. However, when the AI interaction was redesigned with what they call pedagogical friction, so Socratic questioning, guided hints, prompts that forced the student to think before receiving the answer, learning gains were significant. A separate study on Socratic AI tutoring found that students reported significantly greater support for critical, independent, and reflective thinking, directly challenging the narrative that AI necessarily makes you worse at thinking. So based on these studies, the conclusion from the literature is not AI is bad for learning. The conclusion instead is it depends entirely on how you use it. Use it as a substitute for thinking and it atrophies the muscle. Use it as a Socratic partner that pushes you to reason and it can strengthen it. The tool is not the problem. The interaction design is. And right now, nobody's teaching students how to use these tools as thinking partners rather than thinking replacements. Brown's own generative AI and teaching and learning committee, whose report was published the same week as the Srano scandal, found that 56% of undergraduates and 67% of graduate students use generative AI daily or weekly. And here's what I find most striking. Large majorities of those students reported being concerned about the impact on their own learning. They are worried about losing cognitive capacity. They know it's a problem and they're doing it anyway. It's like staring down at a third slice of chocolate cake at 1:00 a.m. Your brain is actively screaming, "This is a terrible choice for a long-term structural integrity." While your hand is already moving for the fork, we need rigorous largescale research, actual control trials with significant participant numbers to understand what's happening here. The Cherikov study is a critical major step in the right direction. We need a lot more like this. We need them across disciplines, across countries, across educational models. We are dealing with a technology that is reshaping how an entire generation interacts with knowledge. And the evidence base for making policy decisions about it is still nowhere near where it should be.
三种毕业生与系统的未来
Serrano的数据中有两个学生值得关注。一个在期中与期末均持续高分,无论有无AI辅助——他真正掌握了材料。Serrano说他很了解这个学生,称其为“优秀学生”。这是你该雇佣的人,是每个雇主都该想要的人——一个真正构建了认知架构的人,能因理解AI运作的领域而将其用作工具,能审视模型输出并判断何时是胡扯。另一个学生持续低分,开卷差、闭卷也差,但他没有作弊。Serrano说:“我钦佩那个人,我理解这种情感。诚实很重要,诚信当然重要。但老实说,钦佩诚实的失败——教育体系走到这一步是非常奇怪的。目标不是高尚的失败,目标是能力。对这个学生正确的反应不是钦佩,而是问如何帮助他们变得更好。”无论有AI还是无AI,无论用不同教学方法还是不同评估模型,不惜一切代价。因为当前体系正在生产三类毕业生:学到了并能证明的人;没学到且能证明的人;没学到但拥有声称他们学到了的文凭的人。第三类应该让每所大学校长夜不能寐,因为市场总会发现真相。当雇主开始发现相当比例的藤校毕业生实际上无法完成学位声称他们能完成的工作时,品牌就会贬值——无论作弊是否被抓住。布朗大学的学位是一个信号,如果信号不再可靠,它就不再有价值。如今招聘数据科学家和技术岗位的方式与传统考试大相径庭——你基本假设他们能够且将会使用AI。老实说,他们为什么不应该呢?评估知识应该在与实际应用时相似的环境中进行。Serrano的学生在其处境下做出了相当理性的反应。而机构的回应——增加监考、取消期中、通过学术准则调查——相当于给有结构问题的建筑配保安:屋顶在塌陷,地基在化为尘土,但谢天谢地我们雇了个拿写字板的人确保没人带未经批准的计算器进来。Cherikov研究呼吁学科层面的评估改革,我同意。但我无法告诉你每个领域具体该怎么做,因为我并非身处每个领域。我能说的是:AI不会消失——我指的不只是大语言模型,而是整个人工智能谱系,工具、系统、基础设施。它已嵌入,正在扩展。每个科学学科都在增长,不断产生更多知识,需要由有限时间和有限寿命的人来吸收。在这个环境中如何教育人,对我们而言已不再是可选项,而是紧迫的。这不是科技公司、监管者或政府能单独回答的问题,它需要我们所有人。无论你从事什么领域、学过什么学科、身处何处,请思考它。思考在你的工作中什么技能真正重要,与你在学校被测试的内容有何不同。思考AI如何与你的领域交汇。请停止感觉AI是发生在你身上的事,开始问如果它是与你一起发生的,会是什么样子。因为在一个足够长的时间线上——我指的是真的很长,比如300年后——我们大多数人都会死。但你们其余人需要留下一个世界,让未来十代人类有能力应对即将到来的挑战。而这始于把教育做对——不是监考、监视或社交媒体上的散点图,而是真正困难的、针对学科的具体工作:弄清楚在一个信息随时可得的世界里学习意味着什么,而真正重要的技能是知道如何运用信息。Serrano教授说“我们不能选择变成白痴”,他是对的。但选择不在于聪明与愚蠢之间,而在于为过去世界设计的教育体系与为现在世界设计的教育体系之间。布朗大学用一个散点图向我们展示了这两者之间的鸿沟有多大。
Original English
But even with what we have, the Brown data tells us something definitely worth sitting with. There are two students in Serrano's data who are worth paying attention to. one scored consistently high on both the midterm and the final with or without AI. They knew the material. Srano said he knew them well and described them as an excellent student. That's the person you hire, right? That's the person every employer should want. Somebody who has actually built the cognitive architecture, who can use AI as a tool because they understand the territory it's operating in, who can look at a model's output and know when it's nonsense. The other students scored consistently low. Poor on the take-home, poor on the final. They didn't cheat. Servano said, I'm quoting, "I admire that person, and I understand the sentiment. Honesty matters. Integrity matters, of course. But I'll be honest with you, admiration for failing honestly is a very strange place for an education system to end up. The goal is not noble failure. The goal is competence. The right response to that student isn't admiration. It's asking how you help them get better." with AI, without AI, with different teaching methods, with different assessment models, whatever it takes. Because right now the system is producing three categories of graduates. People who learned and can prove it, people who didn't learn and can prove it, and people who didn't learn but have a credential that says they did. That third category is the one that should keep every university president awake at night because the market figures it out. It always does. And when employers start discovering that a significant proportion of Ivy League graduates can't actually do the work their degrees claim they can do, the brand degrades whether or not anyone catches the cheating. A degree from Brown is a signal. If the signal stops being reliable, it stops being valuable. Hiring data scientists and technical positions looks very little like a traditional exam these days. You just kind of assume that they can and will use AI. And honestly, why shouldn't they? assessing knowledge should be happening in a similar environment to the same environment they will have when they're actually applying it. Serrano students responded fairly rationally given their circumstances and the institution's response adding proctors avoiding the midterm investigating through the academic code is the educational equivalent of putting a security guard in a building with a structural problem. The roof is skating in. The foundation is turning to dust. But thank God we hired a guy with a clipboard to make sure nobody walks in with an unapproved calculator. The Cherikov study calls for discipline level assessment reform and I agree. But I can't tell you what that looks like for every single field because I don't sit in every field. What I can say is this. AI is not going anywhere. And I don't just mean large language models. I mean the entire spectrum of artificial intelligence, the tools, the systems, the infrastructure. It is embedded. It is expanding. Every scientific discipline is growing, constantly producing more knowledge that needs to be absorbed by people with finite time and finite lifespans. The question of how we educate people in that environment is no longer optional for us. It is urgent. It is not a question that can be answered by tech companies alone or regulators alone or governments alone. It requires all of us. Whatever field you work in, whatever discipline you studied, wherever you sit, please think about it. Think about what skills actually matter in your work versus what you were tested on in school. Think about how AI intersects with your domain. Talk about it here on YouTube, on this channel, with colleagues, with friends, in your own communities. Please stop feeling like AI is something happening to you and start asking what it could look like if it was happening with you. Because on a long enough timeline, and I mean really long here, like 300 years from now, most of us will be dead. Well, except for me because I'm going to be uploaded into the cloud. I'll be hanging out on a server rack in Iceland, thriving on geothermal energy and judging your descendants math skills in real time. But the rest of you will need to have left behind a world where the next 10 generations of humans are equipped to handle what's coming. And that starts with getting education right. Not just proctors or surveillance or scatter plots on social media. I mean the actual difficult discipline specific work of figuring out what it means to learn in a world where the information is always available and the skill that matters is knowing what to do with it. Professor Serrano said we cannot choose to become idiots and he's right. But the choice is not between intelligence and idiocy. It is between an education system designed for the world as it was and one designed for the world as it is. Brown just showed us in one scatterplot how wide the gap between those two things has become.
📌 文中提及的人物和组织
公司/组织: Brown University, Y Combinator, Google DeepMind, MIT Media Lab
产品/模型: ChatGPT