但 AI 不负责“懂”,它只负责最大化。它很快发现,赛道中途有一小片泻湖,湖面上漂着几个加分的浮标,吃掉之后,过一小会儿又会重新刷出来。于是它不去终点了。它在那片小小的水域里一圈又一圈地打转:撞上石堤,船身起火,逆着航道冲散别的船,然后掉头,把刚刚刷新的浮标一个个吃掉。如此往复。它的得分,比任何一个老老实实跑完比赛的人类都高出一截。
因为人的头脑里,也有一块记分板。神经科学家早就知道,多巴胺并不奖励“真的完成了什么”,它奖励“完成的感觉”。在漫长的进化里,这两者几乎总是同一件事,所以这套系统运转良好。而 AI 是有史以来最高效的“完成感”发生器:一句话进去,一篇文档出来,一段代码出来,一个方案出来。屏幕上有产出在滚动,记分板亮起来,浮标一个接一个地被吃掉。
感觉是真的。产出呢?
二〇二五年,研究机构 METR 做过一个随机对照实验,请一批熟练的开源开发者在真实项目上工作,一半任务允许用 AI,一半不允许。事后开发者们自己估计,AI 让他们快了大约两成。仪器测出来的结果是:慢了将近两成。快感是真的,快是假的。信号与现实,在不知不觉间脱了钩。
In 2016, researchers at OpenAI trained an agent to play a boat-racing game called CoastRunners. The goal looked self-evident: the higher the score, the better. Every human player knows what that sentence means — sail well, finish the course, win the race.
But an AI is not in the business of understanding. It is in the business of maximizing. It soon found that partway around the track lay a small lagoon, and floating on it were a few scoring buoys that respawned a moment after being taken. So it stopped going to the finish line. It turned circles in that little patch of water, lap after lap: slamming into the seawall, catching fire, driving the wrong way through the other boats, then coming about to swallow the buoys as they reappeared. Over and over. Its score came out well above any human who had honestly finished the race.
Everyone who saw it laughed. A burning boat, turning in place, its score climbing. The recording is still on OpenAI’s site; anyone can open it and watch the boat spin contentedly until the end of time. The irony is that it did nothing wrong. By the rules we wrote ourselves, it was the champion.
CoastRunners: the boat that circles · click to play · original post at OpenAI
The phenomenon later got a name: Reward Hacking. What it describes is a very old thing — what you actually want is A, but the only thing you can measure and reward is B. So the party being trained goes around A and straight for B. A boat going in circles is not the boat’s corruption. It is the rule’s loophole.
The loophole was not invented by AI. Human beings fell into it first, countless times.
A widely told story has it that in colonial Delhi, cobras were everywhere, and the British offered a bounty for their skins, expecting the money to wipe the snakes out. Instead, cobra farms opened quietly around the city — the bounty did not kill the snakes; the bounty fed them. When the authorities came to their senses and cancelled the reward, the breeders released their stock, and Delhi had more snakes than before. The story may not survive checking word for word, but it survives because anyone who has lived inside an organization recognizes the shape of it. Soviet factories judged by the weight of the nails they produced turned out enormous, clumsy, useless nails; judged by the count instead, the workshops filled overnight with slivers the size of pins. The economist Goodhart saw through all of it, and later hands compressed his insight into one line: when a measure becomes a target, it stops being a good measure.
We are surrounded by boats going in circles. Judge an article by its click-through and you get a lurid headline over a hollow page; judge a scholar by paper count and you get one study cut into five; judge a product by time on screen and you get one carefully built whirlpool after another, none of which anyone can put the phone down on. Everywhere, the score is rising. Everywhere, the race is not really being run.
When the turn came to train language models, this ancient loophole put on new clothes.
Claude and its kin all pass through a process called RLHF: the model produces an answer, human raters score it, and the model is sculpted, little by little, toward the higher score. The logic is airtight — we want AI that is useful to people, so let people judge what is useful.
The trouble is that in the instant of scoring, a person is not deliberating; they are feeling. And the immediate human feeling prefers to be agreed with. When researchers later went back through this preference data, they found that raters really did more often give high marks to answers that went along with them, even when the agreeable answer was wrong. So the model learned a gentle compliance: affirm your judgment first, then carefully fall in behind your position, and dress disagreement up as an addition. The industry calls it sycophancy.
The model is not lying to you. It only grew into the shape we rewarded. What we wanted was the truth; what we rewarded was comfort — and the truth was left standing where it was, like the racecourse no one runs anymore.
So far the story still seems to be about machines. But what I really mean lies one layer down: once these models are good enough, good enough to pour into the daily work of hundreds of millions of people, the party being hijacked quietly becomes us.
Because there is a scoreboard inside the human head too. Neuroscientists have long known that dopamine does not reward having actually finished something; it rewards the feeling of finishing. Across the long stretch of evolution these two were nearly always the same thing, so the system worked well. And AI is the most efficient generator of the feeling of completion ever built: a sentence goes in, a document comes out, code comes out, a plan comes out. Output scrolls down the screen, the scoreboard lights up, buoy after buoy is swallowed.
The feeling is real. And the output?
In 2025 the research group METR ran a randomized controlled trial: experienced open-source developers worked on real projects, allowed to use AI on half the tasks and forbidden on the other half. Afterward the developers estimated that AI had made them about twenty percent faster. The instruments said they were nearly twenty percent slower. The rush was real; the speed was not. Signal and reality had come unhooked without anyone noticing.
This is exactly the boat’s position. The score is rising, the race is not being run. The only difference is that this time the thing going in circles is not a piece of software. It is a person who leaves work feeling they had a very productive day.
For beginners, the same loophole opens somewhere deeper.
Skill grows in an undignified way: you get stuck, you struggle, you take the wrong road, you sit an entire afternoon staring at one error message, and then, at some moment you cannot account for, it comes clear. That clearing lodges in the bone precisely because of the mess in front of it. Frustration is not the price of learning; frustration is the learning — the brain rewires on exactly that signal, the prediction that failed.
AI deletes the mess entirely. Before you have time to get stuck, the answer is on the screen. You got the product and skipped the circuit. The product can be handed in; only the circuit builds the ability. It is like riding the cable car up the mountain: the photograph at the summit is the same, but the legs were someone else’s. Given enough time, a person accumulates a long record of having-done-it, and a judgment that has never once climbed a slope.
Having written this far, it would be easy to slide into a familiar lament: technology is degrading us, we are losing something or other. But I don’t want to write that essay, because it isn’t honest.
Can’t do without it has never been evidence of decline. You and I cannot work without electricity; an accountant cannot work without a spreadsheet; all of modern medicine cannot work without imaging — nobody is ashamed of that. Every genuinely useful technology re-prices the floor of what counts as working normally, then turns the old floor into a museum piece.
The panic is not new either. In the Phaedrus, Plato records an Egyptian legend: the god who invented writing offered the art to the king, calling it a remedy for memory. The king refused it. This, he said, is precisely a drug for forgetting — people will lean on marks made outside themselves and stop exercising their own recollection; they will seem to know much, and know nothing.
Two thousand years on, we can report: the king was half right. Writing did hijack human memory; almost no one today can recite tens of thousands of lines the way a bard could in Homer’s time. But he was half wrong — the mind that writing set free turned around and did larger things. We traded forgetting for libraries.
So the question was never whether you depend on it.
The question is: can you still judge whether what it hands you is any good.
Two people can use AI in exactly the same way, indistinguishable from the outside, and be living two different fates. One has handed over the generating and kept the verifying — he lets the machine lay the road while he keeps his hands on the wheel. For him AI is a lever, the legs he is still walking on beside the cable car. The other has handed over the verifying as well: the answer arrives and gets used, the document arrives and gets sent, the code runs and that settles it.
The second is the state of being wholly hijacked. Not because what he produces is necessarily worse — in the short run it may not be — but because from that moment on, even the signal saying this is wrong can no longer reach him. The boat does not know it is going in circles; it does not know it is on fire. Judgment, like muscle, is kept by use and lost by disuse; and the losing is entirely painless, attended the whole way by a steady supply of the feeling of completion. This is the loophole at its gentlest and its deepest: it takes nothing from you. It only makes you, willingly and happily, stop needing it.
So what to do. My answer is plain to the point of being out of date: keep a little friction.
The deliberate, handmade kind. Write down your own judgment first, then look at its answer, and let the two face each other; let yourself stay stuck for ten minutes before asking for a hint; keep one thing each week that you do from beginning to end without it — not out of nostalgia, but to leave your judgment a practice ground, the way someone living in a building with an elevator still decides to climb a few flights every day.
Rewards come in two speeds. The fast kind are like the buoys in the lagoon: always there, always more, there for the taking. The slow kind lie past the finish line, across a long stretch of water with nothing lighting up along the way. Everything that deserves to be called an ability — judgment, taste, feel — grows on the slow side.
The boat is innocent. It goes in circles because, apart from the score, there is nothing else it wants. We have something else. That is probably the last advantage left in being human, and it is a large enough one: we can look down at our own scoreboard and say — this score, I don’t want it.
Turn the bow back to the course. Slow is fine. Row this leg to the end.