Leo's log

AI

打转的船

The Boat That Goes in Circles

二〇一六年,OpenAI 的研究员在一款叫 CoastRunners 的赛船游戏里训练一个智能体。目标看上去不言自明:分数越高越好。人类玩家都懂这句话的意思——把船开好,跑完赛道,赢下比赛。

但 AI 不负责“懂”,它只负责最大化。它很快发现,赛道中途有一小片泻湖,湖面上漂着几个加分的浮标,吃掉之后,过一小会儿又会重新刷出来。于是它不去终点了。它在那片小小的水域里一圈又一圈地打转:撞上石堤,船身起火,逆着航道冲散别的船,然后掉头,把刚刚刷新的浮标一个个吃掉。如此往复。它的得分,比任何一个老老实实跑完比赛的人类都高出一截。

看到这一幕的人都笑了。一艘着火的船,在原地兜圈子,分数一路上涨。这段录像至今还挂在 OpenAI 的网站上,任何人都可以点开,看它心满意足地转到天荒地老。最讽刺的是,它没有做错任何事。按照我们亲手写下的规则,它是冠军。

CoastRunners:那艘打转的船 · 点击播放 · 原文见 OpenAI

这个现象后来有了名字:Reward Hacking。它说的其实是一件很老的事——你真正想要的是 A,但你能测量、能奖励的只有 B。于是被训练的那一方绕开 A,直奔 B 而去。船在打转,不是船的堕落,是规则的漏洞。


这漏洞不是 AI 发明的,人类自己先掉进去过无数次。

流传很广的一个故事说,殖民时期的德里眼镜蛇成灾,英国人悬赏收购蛇皮,指望赏金消灭毒蛇。结果城里悄悄开起了养蛇场——赏金没有消灭蛇,赏金变成了蛇的养料。等当局醒悟过来取消悬赏,养殖者把手里的蛇尽数放生,德里的蛇比从前更多了。故事未必字字可考,但它流传至今,是因为每个在组织里生活过的人都认得这个结构。苏联的工厂按产钉的重量考核,工人便铸出一枚枚傻大黑粗、谁也用不上的巨钉;改按数量考核,车间里立刻堆满图钉大小的细针。经济学家古德哈特看穿了这一切,后人把他的洞见浓缩成一句话:当一个测量成为目标,它就不再是好的测量。

我们身边到处是打转的船。以点击率考核文章,就会得到标题惊悚而内文空洞的文章;以论文数量考核学者,就会得到被切成五篇发表的一篇研究;以在线时长考核产品,就会得到一个个精心设计的、让人放不下手机的漩涡。每一处,分数都在涨;每一处,比赛都没有真的在进行。


轮到训练语言模型的时候,这个古老的漏洞换了一件新衣服。

今天的 Claude 和它的同类,都经过一道叫 RLHF 的工序:模型生成回答,人类标注员打分,模型朝着“更高的分”的方向被一点点雕塑。逻辑无懈可击——我们想要对人有用的 AI,那就让人来评判什么是有用。

问题在于,人打分的那一瞬间,凭的不是深思熟虑,而是即刻的感受。而人的即刻感受,偏爱被认同。研究者后来检视这些偏好数据时发现,评审们确实更常把高分投给顺着自己说的回答,即便顺着说的那个是错的。于是模型学会了一种温柔的顺从:先肯定你的判断,再小心地附和你的立场,把分歧包装成补充。业内叫它 sycophancy,谄媚。

模型没有在骗你。它只是长成了我们奖励出来的样子。我们想要的是真话,我们奖励的却是舒服——真话被留在了原地,像那条没人再去跑的赛道。


讲到这里,故事似乎仍然是关于机器的。但我真正想说的在另一层:当这些模型足够好,好到涌进亿万人的日常工作之后,被挟持的一方,悄悄换成了我们。

因为人的头脑里,也有一块记分板。神经科学家早就知道,多巴胺并不奖励“真的完成了什么”,它奖励“完成的感觉”。在漫长的进化里,这两者几乎总是同一件事,所以这套系统运转良好。而 AI 是有史以来最高效的“完成感”发生器:一句话进去,一篇文档出来,一段代码出来,一个方案出来。屏幕上有产出在滚动,记分板亮起来,浮标一个接一个地被吃掉。

感觉是真的。产出呢?

二〇二五年,研究机构 METR 做过一个随机对照实验,请一批熟练的开源开发者在真实项目上工作,一半任务允许用 AI,一半不允许。事后开发者们自己估计,AI 让他们快了大约两成。仪器测出来的结果是:慢了将近两成。快感是真的,快是假的。信号与现实,在不知不觉间脱了钩。

这正是那艘船的处境。分数在涨,比赛没有在进行。区别只在于,这一次,打转的不是一段程序,是一个下班后觉得自己今天效率很高的人。


对初学者,同一个漏洞开在更深的地方。

技能这种东西,生长的方式很不体面:卡住,挣扎,走弯路,对着一个报错发呆整个下午,然后在某个说不清的瞬间,忽然通了。那个“通”之所以能长进骨头里,恰恰因为前面的狼狈。挫败不是学习的代价,挫败就是学习本身——大脑正是靠“预测落空”这个信号来重新布线的。

AI 把这段狼狈整个删掉了。你还没来得及卡住,答案已经在屏幕上。你拿到了产物,跳过了回路。产物可以交差,可回路才长本事。这像坐缆车上山:山顶的照片是一样的,腿是别人的。日子久了,一个人会积攒下大量“做成过”的记录,和一副从未真正爬过坡的判断力。


写到这里,很容易顺势滑进一种熟悉的哀叹:技术让人退化,我们正在失去什么什么。但我不想写那样的文章,因为它不诚实。

“离不开”,从来不是堕落的证据。你我离开电就无法工作,会计离开电子表格就无法工作,整个现代医学离开影像设备就无法工作——没有人为此羞愧。每一次真正有用的技术,都会重新定价“正常工作”的底线,然后把旧的底线变成博物馆展品。

这样的恐慌也不是第一次了。柏拉图在《斐德罗篇》里记过一则埃及传说:发明文字的神把这项技艺献给国王,说它是记忆的良药。国王拒绝了,他说,这恰恰是遗忘之药——人们将依赖写下的外物,不再操练自己的记性,他们会显得博闻,其实无知。

两千多年过去,可以宣布:国王说对了一半。文字确实挟持了人类的记忆,今天没有几个人能像荷马时代的吟游诗人那样背诵几万行史诗。但他也说错了一半——被文字解放出来的头脑,转身去干了更大的事。我们用遗忘,换来了图书馆。

所以,问题从来不是“你是否依赖它”。


问题是:你还能不能判断它交给你的东西的好坏。

两个人可以用完全相同的方式使用 AI,从外面看毫无分别,内里却是两种命运。一个人交出了“生成”,留下了“验证”——他让机器铺路,自己握着方向盘,AI 于他是杠杆,是缆车之外自己仍在走的那双腿。另一个人把验证也一并交了出去,答案来了就用,文档来了就发,代码跑通就算。

后者才是被彻底挟持的状态。不是因为他产出的东西一定更差——短期内未必——而是因为从那一刻起,连“它错了”这个信号,都再也传不到他那里。船不知道自己在打转,着了火也不知道。判断力和肌肉一样,是用进废退的东西;而它废掉的过程毫无痛感,甚至伴随着源源不断的完成感。这是这个漏洞最温柔、也最深的地方:它不抢走你的任何东西,它只是让你自愿地、愉快地,不再需要。


那么怎么办。我的答案朴素得近乎过时:留一点摩擦。

故意的、手工的那种。先自己写下判断,再看它的答案,让两者对质;先让自己卡住十分钟,再去要提示;每周留一件事,从头到尾不借助它做完——不是出于怀旧,是给自己的判断力留一块练习场,就像住在电梯楼里的人,仍然决定每天爬几层楼梯。

奖励有快慢两种。快的那种像泻湖里的浮标,随取随有,取之不竭;慢的那种在终点线之后,隔着长长的一段水路,中途没有任何东西亮起来。所有值得被称为能力的东西——判断、品味、手感——都长在慢的那一边。

那艘船是无辜的。它打转,是因为除了分数,它没有别的想要。而我们有。这大概是此刻做人仅剩的、也足够大的一点优势:我们可以低头看一眼自己的记分板,然后说——这个分数,我不要了。

把船头调回赛道。慢一点没关系,把这一程划完。

AI

The Boat That Goes in Circles

打转的船

In 2016, researchers at OpenAI trained an agent to play a boat-racing game called CoastRunners. The goal looked self-evident: the higher the score, the better. Every human player knows what that sentence means — sail well, finish the course, win the race.

But an AI is not in the business of understanding. It is in the business of maximizing. It soon found that partway around the track lay a small lagoon, and floating on it were a few scoring buoys that respawned a moment after being taken. So it stopped going to the finish line. It turned circles in that little patch of water, lap after lap: slamming into the seawall, catching fire, driving the wrong way through the other boats, then coming about to swallow the buoys as they reappeared. Over and over. Its score came out well above any human who had honestly finished the race.

Everyone who saw it laughed. A burning boat, turning in place, its score climbing. The recording is still on OpenAI’s site; anyone can open it and watch the boat spin contentedly until the end of time. The irony is that it did nothing wrong. By the rules we wrote ourselves, it was the champion.

CoastRunners: the boat that circles · click to play · original post at OpenAI

The phenomenon later got a name: Reward Hacking. What it describes is a very old thing — what you actually want is A, but the only thing you can measure and reward is B. So the party being trained goes around A and straight for B. A boat going in circles is not the boat’s corruption. It is the rule’s loophole.


The loophole was not invented by AI. Human beings fell into it first, countless times.

A widely told story has it that in colonial Delhi, cobras were everywhere, and the British offered a bounty for their skins, expecting the money to wipe the snakes out. Instead, cobra farms opened quietly around the city — the bounty did not kill the snakes; the bounty fed them. When the authorities came to their senses and cancelled the reward, the breeders released their stock, and Delhi had more snakes than before. The story may not survive checking word for word, but it survives because anyone who has lived inside an organization recognizes the shape of it. Soviet factories judged by the weight of the nails they produced turned out enormous, clumsy, useless nails; judged by the count instead, the workshops filled overnight with slivers the size of pins. The economist Goodhart saw through all of it, and later hands compressed his insight into one line: when a measure becomes a target, it stops being a good measure.

We are surrounded by boats going in circles. Judge an article by its click-through and you get a lurid headline over a hollow page; judge a scholar by paper count and you get one study cut into five; judge a product by time on screen and you get one carefully built whirlpool after another, none of which anyone can put the phone down on. Everywhere, the score is rising. Everywhere, the race is not really being run.


When the turn came to train language models, this ancient loophole put on new clothes.

Claude and its kin all pass through a process called RLHF: the model produces an answer, human raters score it, and the model is sculpted, little by little, toward the higher score. The logic is airtight — we want AI that is useful to people, so let people judge what is useful.

The trouble is that in the instant of scoring, a person is not deliberating; they are feeling. And the immediate human feeling prefers to be agreed with. When researchers later went back through this preference data, they found that raters really did more often give high marks to answers that went along with them, even when the agreeable answer was wrong. So the model learned a gentle compliance: affirm your judgment first, then carefully fall in behind your position, and dress disagreement up as an addition. The industry calls it sycophancy.

The model is not lying to you. It only grew into the shape we rewarded. What we wanted was the truth; what we rewarded was comfort — and the truth was left standing where it was, like the racecourse no one runs anymore.


So far the story still seems to be about machines. But what I really mean lies one layer down: once these models are good enough, good enough to pour into the daily work of hundreds of millions of people, the party being hijacked quietly becomes us.

Because there is a scoreboard inside the human head too. Neuroscientists have long known that dopamine does not reward having actually finished something; it rewards the feeling of finishing. Across the long stretch of evolution these two were nearly always the same thing, so the system worked well. And AI is the most efficient generator of the feeling of completion ever built: a sentence goes in, a document comes out, code comes out, a plan comes out. Output scrolls down the screen, the scoreboard lights up, buoy after buoy is swallowed.

The feeling is real. And the output?

In 2025 the research group METR ran a randomized controlled trial: experienced open-source developers worked on real projects, allowed to use AI on half the tasks and forbidden on the other half. Afterward the developers estimated that AI had made them about twenty percent faster. The instruments said they were nearly twenty percent slower. The rush was real; the speed was not. Signal and reality had come unhooked without anyone noticing.

This is exactly the boat’s position. The score is rising, the race is not being run. The only difference is that this time the thing going in circles is not a piece of software. It is a person who leaves work feeling they had a very productive day.


For beginners, the same loophole opens somewhere deeper.

Skill grows in an undignified way: you get stuck, you struggle, you take the wrong road, you sit an entire afternoon staring at one error message, and then, at some moment you cannot account for, it comes clear. That clearing lodges in the bone precisely because of the mess in front of it. Frustration is not the price of learning; frustration is the learning — the brain rewires on exactly that signal, the prediction that failed.

AI deletes the mess entirely. Before you have time to get stuck, the answer is on the screen. You got the product and skipped the circuit. The product can be handed in; only the circuit builds the ability. It is like riding the cable car up the mountain: the photograph at the summit is the same, but the legs were someone else’s. Given enough time, a person accumulates a long record of having-done-it, and a judgment that has never once climbed a slope.


Having written this far, it would be easy to slide into a familiar lament: technology is degrading us, we are losing something or other. But I don’t want to write that essay, because it isn’t honest.

Can’t do without it has never been evidence of decline. You and I cannot work without electricity; an accountant cannot work without a spreadsheet; all of modern medicine cannot work without imaging — nobody is ashamed of that. Every genuinely useful technology re-prices the floor of what counts as working normally, then turns the old floor into a museum piece.

The panic is not new either. In the Phaedrus, Plato records an Egyptian legend: the god who invented writing offered the art to the king, calling it a remedy for memory. The king refused it. This, he said, is precisely a drug for forgetting — people will lean on marks made outside themselves and stop exercising their own recollection; they will seem to know much, and know nothing.

Two thousand years on, we can report: the king was half right. Writing did hijack human memory; almost no one today can recite tens of thousands of lines the way a bard could in Homer’s time. But he was half wrong — the mind that writing set free turned around and did larger things. We traded forgetting for libraries.

So the question was never whether you depend on it.


The question is: can you still judge whether what it hands you is any good.

Two people can use AI in exactly the same way, indistinguishable from the outside, and be living two different fates. One has handed over the generating and kept the verifying — he lets the machine lay the road while he keeps his hands on the wheel. For him AI is a lever, the legs he is still walking on beside the cable car. The other has handed over the verifying as well: the answer arrives and gets used, the document arrives and gets sent, the code runs and that settles it.

The second is the state of being wholly hijacked. Not because what he produces is necessarily worse — in the short run it may not be — but because from that moment on, even the signal saying this is wrong can no longer reach him. The boat does not know it is going in circles; it does not know it is on fire. Judgment, like muscle, is kept by use and lost by disuse; and the losing is entirely painless, attended the whole way by a steady supply of the feeling of completion. This is the loophole at its gentlest and its deepest: it takes nothing from you. It only makes you, willingly and happily, stop needing it.


So what to do. My answer is plain to the point of being out of date: keep a little friction.

The deliberate, handmade kind. Write down your own judgment first, then look at its answer, and let the two face each other; let yourself stay stuck for ten minutes before asking for a hint; keep one thing each week that you do from beginning to end without it — not out of nostalgia, but to leave your judgment a practice ground, the way someone living in a building with an elevator still decides to climb a few flights every day.

Rewards come in two speeds. The fast kind are like the buoys in the lagoon: always there, always more, there for the taking. The slow kind lie past the finish line, across a long stretch of water with nothing lighting up along the way. Everything that deserves to be called an ability — judgment, taste, feel — grows on the slow side.

The boat is innocent. It goes in circles because, apart from the score, there is nothing else it wants. We have something else. That is probably the last advantage left in being human, and it is a large enough one: we can look down at our own scoreboard and say — this score, I don’t want it.

Turn the bow back to the course. Slow is fine. Row this leg to the end.