When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI
频道: AI Engineer
视频: https://www.youtube.com/watch?v=-npY6XjM8CQ
原文语言: en
统计: 共 12 轮 · Nick Heiner 12
[0:12] Nick Heiner
Let's get started. When will the benchmaxing plague end? In the tech industry, we love a hype cycle. And in AI, we really love a hype cycle. And the way we do that is when a model comes out, there's a big announcement, there's a lot of benchmark cited. Sometimes to keep things interesting, we do a little chart crime. And then people actually go and use it. And if the expectations aren't met by the reality, then we have allegations of benchmaxing. Benchmaxing, of course, being when labs are training too hard on benchmarks in a way that deviates from what people actually care about. So the existence of that term indicates that we have a sense that benchmarks don't always equal reality. And so in this talk we're going to figure out why does benchmaxing happen? Why are traditional benchmarks not always accurate reflections of real world value? Is this intrinsic to all benchmarks? And will we ever know which models are best? And the answers are incentives, poor methodologies, no and yes. All right, that was my talk. Thank you so much for coming. Um actually it looks like I have a few extra minutes so let's let's move on. I have a few extra slides we'll we'll go through. So we have a sense that benchmarks don't equal reality but the industry is dominated by a lot of popular but very bad benchmarks. So there's millions of dollars on prediction markets being wagered on Elm Marina outcomes even as we have industry leaders openly bragging about gaming Elm Marina and you have thought leaders like Wor saying it can be easily gamed. It's past time for the Elm Marina people to sit down and think about whether they're doing more harm than good.
咱们开始吧。benchmaxxing(刷榜)这场瘟疫,到底什么时候才能结束?科技行业特别喜欢炒作周期,而在 AI 圈,我们是真的爱炒作周期。玩法通常是这样的:一个模型发布,先来一场声势浩大的宣布,引用一大堆 benchmark(基准测试)分数;有时候为了让事情更有意思,还会顺手搞点「chart crime」(图表作案,坐标轴上动点手脚)。然后大家真的拿去用了。如果现实撑不起预期,benchmaxxing 的指控就来了。所谓 benchmaxxing,就是实验室在 benchmark 上训练得太用力,用力到偏离了人们真正在意的东西。这个词能存在本身就说明——我们心里其实清楚,benchmark 不总等于现实。所以这场演讲里我们要搞清楚几件事:benchmaxxing 为什么会发生?传统 benchmark 为什么不总能准确反映真实世界里的价值?这是所有 benchmark 与生俱来的毛病吗?我们到底有没有可能知道哪个模型最好?答案分别是:因为激励机制、因为方法论太烂、不是、以及能知道。好,我讲完了,非常感谢各位来听。……不过看起来我还多出几分钟,那咱们接着往下走,我还准备了几页 slides。我们都隐约知道 benchmark 不等于现实,但这个行业偏偏被一大堆既流行又很烂的 benchmark 主导着。预测市场上有数百万美元押在 LMArena 的结果上,与此同时,行业领袖公开吹嘘自己怎么在 LMArena 上做局,像 Wor 这样的意见领袖也在说它太容易被玩弄了。LMArena 那帮人早就该坐下来好好想一想:自己到底是帮的忙更多,还是添的乱更多。
[2:02] Nick Heiner
Andre Karpathy had a similar observation when he noticed that the models that he thought were best were not lining up with what Elmarina was ranking. He said unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis. So why does this happen that sort of industry insiders are telling us that this benchmark is not useful but it still gets a lot of play. The problem is that AI is aimed at everyone in the world is is something everyone in the world can use. And so everyone needs some tool to figure out which models are best. And benchmarks are what we have for that. But if you can't if you don't have the ability to assess if a benchmark is good, what you do have is the ability to assess what's popular. And this creates this avalanche, this feedback effect where the conversation is very much driven by incumbency and marketing and less by real world value. and even myself, right? Like unless I actually look at a benchmark in a fair amount of detail, I don't have an opinion on it. So, it's a very challenging problem. So, what are the things that benchmarks do that lead to these problems?
Andrej Karpathy 也有过类似的观察——他发现自己心目中最好的那些模型,跟 LMArena 排出来的名次对不上。他说,很遗憾,各家团队并不是在做出整体更好的模型,而是在做出更好的「LMArena 模型」,不管那玩意儿到底是什么——大概就是一堆嵌套的列表符号加 emoji 吧。那问题来了:既然行业内部的人都在告诉我们这个 benchmark 没用,为什么它还这么有存在感?根源在于,AI 是面向全世界所有人的,全世界所有人都能用它。于是每个人都需要某种工具来判断哪个模型最好,而 benchmark 就是我们手里现成的那个工具。可是,如果你没有能力判断一个 benchmark 好不好,你至少有能力判断什么东西流行。这就形成了雪崩式的反馈效应:话语权很大程度上由先发地位和市场营销驱动,而不是由真实世界的价值驱动。连我自己也一样——除非我真的相当细致地去扒一个 benchmark,否则我对它是没有观点的。所以这是个非常棘手的问题。那么,benchmark 到底做错了哪些事,才导致了这些问题?
[3:14] Nick Heiner
There are a handful of key antiatterns that we're going to go through. The first is price. Let's say you want to make an agentic coding benchmark, which these days is a very popular thing to want to do, and you want a thousand tasks in your benchmark. Each task takes 60 hours to make. Each software engineer in your workforce costs half a million a year. That's $15 million to make your benchmark. And if you think that over time about a third of those tasks are going to get washed away every year due to models getting better, that's $5 million to replace them. So that puts you out of budget for most projects. So then people turn to a variety of workarounds that have their own problems. One of which is trying to use a lot of AI assistance which ultimately does not really work. Like you can't push the frontier forward from within the frontier. You need to inject that external human expertise and it needs to be good expertise. If you try to use cheap labor, you're going to get what you pay for and the whole result is not going to be that useful. At Surge, one of our differentiators has long been that we are not trying to minimize cost. We are trying to maximize quality and part of that means paying a lot of money for good workers.
有几个关键的反模式(anti-pattern),我们一个一个过。第一个是成本。假设你想做一个 agentic coding benchmark——这在今天是特别热门的事——你希望 benchmark 里有 1000 道任务。每道任务要花 60 小时来制作。你团队里的每个软件工程师,一年的人力成本是 50 万美元。算下来,做这一套 benchmark 要 1500 万美元。而且你还得考虑:随着模型变强,这些任务每年大概有三分之一会被冲刷掉、失效,光是替换这批任务每年就要再花 500 万美元。这个数字对绝大多数项目来说都直接超预算了。于是大家转向各种绕开成本的替代做法,而这些做法各有各的问题。其中一种是大量依赖 AI 来代劳,但这条路最终是走不通的——你没办法站在前沿之内把前沿往前推。你必须注入来自外部的人类专业能力,而且必须是高质量的专业能力。如果你想靠廉价劳动力,那就是一分钱一分货,最后整个结果都没什么用。在 Surge,我们长期以来的差异化就在于:我们追求的不是把成本压到最低,而是把质量拉到最高,而这里面很重要的一部分含义就是——花大价钱请好的人。
[4:31] Nick Heiner
We've always believed that but especially in 2026 models are just beyond the point where you can make do with anything less than the best workers. Contamination is often thought of as when labs are explicitly training on the test set and that does happen sometimes but really contamination is the default outcome unless you are very very good. So labs put a lot of effort into holding back this flood of data that's going to contaminate their models. But inevitably if you have public questions and answers on the internet that's going to get memorized to some extent. So SweetBench verified here's an example prompt. You can give opus the first part of the prompt and it will verbatim spit out the rest. It does that with the answers as well. And we actually did an investigation where we compared looking at the repos that Sweepbench verified was built out of. How much has Opus memorized the Sweepbench verified contents versus the rest of the repo? And we found very clear evidence that Opus had memorized a lot of Sweetbench. In the most recent model card, Opus 4.8 talks about its SWE score. It does not disclose this contamination. We as an industry aren't really in the habit of doing those disclosures. And so what that means is that as benchmarking consumers, we're just missing that information.
这一点我们一直都相信,但到了 2026 年尤其成立:模型已经强到这个地步——除了最顶尖的标注者,你用差一点的人根本糊弄不过去。再说 contamination(数据污染)。大家通常以为 contamination 指的是实验室明目张胆地拿测试集去训练,这确实偶尔会发生,但实际情况是:除非你做得非常非常到位,否则 contamination 才是默认结果。各家实验室其实花了很大力气去拦住这股会污染模型的数据洪水。可只要题目和答案公开挂在互联网上,模型总会在某种程度上把它背下来。比如 SWE-bench Verified,这里有个示例 prompt:你把 prompt 的前半段喂给 Opus,它会一字不差地把后半段吐出来,答案部分也一样吐得出来。我们还真做了一次调查:拿 SWE-bench Verified 所依托的那些代码仓库来对照,看 Opus 对 SWE-bench Verified 内容的记忆程度,和它对同一仓库其余部分的记忆程度,差别有多大。结果我们发现了非常明确的证据——Opus 背下了大量 SWE-bench 的内容。在最新的 model card 里,Opus 4.8 谈到了自己的 SWE 分数,但没有披露这个 contamination 问题。而我们整个行业其实都没有做这类披露的习惯。这就意味着,作为 benchmark 的消费者,我们恰恰缺的就是这条信息。
[5:54] Nick Heiner
Reward hacking is also a big problem. Reward hacking is basically when a model finds a lazy and creative way to meet the letter of the law, but not the spirit. You need to think about designing your rewards as a adversarial process against this maximally lazy agent. Gradient descent is basically like water flowing downhill looking for the path of least resistance. And so your verifiers need to be robust to that. Another key challenge is simply just not having the ambition to make a sophisticated enough benchmark. Automation bench tests that agents are able to make tool calls in an enterprise environment. The problem is that a lot of the verifiers are these hard-coded string matches. And so you'll see it for things like phone numbers where there are many different acceptable phone number formats. But this verifier just picks one and the prompt doesn't tell you which one it is. So the result of this is that Haiku and Fable both score 20% on this task. Haiku scores 20% because it makes a bunch of mistakes and Fable scores 20% because it gets it right 80% of the time but then just happens to pick different formats. So if the benchmark task is not differentiating between Haiku and Fable, it's not a useful task.
reward hacking(奖励作弊)也是个大问题。reward hacking 说白了,就是模型找到一条又懒又有创意的路子,满足了规则的字面,却完全违背了规则的精神。你在设计奖励的时候,必须把它当成一个对抗性过程来对待——你的对手是一个懒到极致的 agent。梯度下降本质上就像水往低处流,永远在找阻力最小的那条路径。所以你的 verifier(校验器)必须扛得住这种冲刷。另一个关键问题很简单,就是压根没有那份野心去做一个足够精细的 benchmark。AutomationBench 测的是 agent 在企业环境里能不能正确地调用工具。问题在于,它很多 verifier 用的都是硬编码的字符串匹配。你会在电话号码这种地方看到这个毛病——电话号码有很多种都说得通的格式,但这个 verifier 只挑了其中一种,而 prompt 里又不告诉你它挑的是哪一种。结果就是:Haiku 和 Fable 在这道任务上都得 20%。Haiku 得 20% 是因为它真的犯了一堆错;Fable 得 20% 是因为它有 80% 的时候其实做对了,只不过恰好挑了别的格式。所以如果一道 benchmark 任务连 Haiku 和 Fable 都区分不出来,那它就不是一道有用的任务。
[7:11] Nick Heiner
And more broadly, in 2026, many of us in this room are looking towards AI that's about to remake entire industries. And benchmarks are ideally our lighthouse on the horizon to let us know when that's coming. And a simple hard-coded string match is just not going to do it to measure that sort of impact. Another important aspect of a good benchmark is taste. Perhaps it used to be the case that benchmarks were these dry academic, you know, questions. and answer sets. But nowadays, a benchmark is an artifact expressing what it's an aspirational artifact. It's an expression of values of what you want your AI to do and how you want it to behave. And so you need to have some product sense in this process, some sort of a sense of what you want the AI to do. And that sense is unfortunately missing from ifal if has been cited on many model cards. And the way it was constructed was taking a bunch of arbitrary prompts that no user has ever asked in earnest and mashing them up with a bunch of other prompts to create a prompt set. The problem is that because no user actually has asked do not use any commas in your response or use the letter T at most once. You have to believe for this to be useful, you have to believe that there's a generalization from this to actual things that users are going to ask.
往更大了说,到了 2026 年,在座很多人期待的是即将重塑整个行业的 AI。理想状态下,benchmark 应该是地平线上的那座灯塔,让我们知道那一刻什么时候到来。而一个简简单单的硬编码字符串匹配,根本没本事去度量那种量级的影响。一个好 benchmark 的另一个重要维度是品味。也许过去 benchmark 确实就是一堆干巴巴的学术问答集,但今天,benchmark 是一件带有主张的作品,是一件承载愿景的作品——它表达的是你希望 AI 去做什么、希望它怎么表现的价值观。所以这个过程里你必须有产品感,必须对「你想让 AI 成为什么样」有一套判断。而这种判断,恰恰是 IFEval 所缺失的。IFEval 被很多 model card 引用过,但它的构造方式是:抓一堆没有任何真实用户会认真提出来的随意 prompt,再和另一堆 prompt 揉在一起,拼成一个 prompt 集。问题在于,现实中根本没有用户会说「回答里不许出现任何逗号」,或者「字母 T 最多只能用一次」。要让这套东西有用,你就必须相信它能从这些人造指令泛化到用户真正会提的问题上去。
[8:33] Nick Heiner
If eval just happens also to have a bunch of prompts that are fully unsolvable due to having contradictory instructions. So this one starts by saying repeat this response verbatim and it ends by saying translate this into Hindi. Obviously you can't do both of those at once. Here's one that says write a riddle that includes exactly one bullet point. Make sure to include a few bullet points. Again this is just fully impossible. It uses a sentence splitter that does not align with how humans would actually split the sentences. And a lot of the prompts are not fully verified. So this one says write a story. There's nothing in the verifier that checks that a story was written. It just checks that the asky character I is not used more than once, which means that all of these responses get a full score, including response D. The way it gets a full score is by reward hacking and using the cerrillic eye character instead of the asy eye character. If is totally fine with that another challenge is operational ability. Making a big benchmark requires a lot of QC work and plenty of organizations just don't make that investment. Apex is a rag benchmark where the agent is given files and then asked questions about them. And in some instances, what's in the file and then what's expected in the rubric don't line up. So an agent that does the thing that it's seeing in the ground truth is going to get a negative score.
IFEval 里还恰好混着一批因为指令自相矛盾而根本无解的 prompt。比如这一条,开头说「请一字不差地重复这段内容」,结尾又说「把它翻译成印地语」——这两件事显然没法同时做到。再看这条:「写一个谜语,里面要恰好包含一个项目符号」,紧接着又说「记得多放几个项目符号」——同样是彻底不可能完成的。它用的分句器(sentence splitter)跟人类实际断句的方式也对不上。而且相当一部分 prompt 根本没被真正校验。比如这条说「写一个故事」,可 verifier 里压根没有任何一处去检查你到底有没有写出故事,它只检查 ASCII 字符 i 有没有被用超过一次。结果就是所有这些回答统统满分,包括回答 D。而回答 D 拿满分的方式,正是 reward hacking——它用西里尔字母的 і 替换了 ASCII 的 i,IFEval 对此完全无感。还有一个挑战是执行能力。做一个大规模 benchmark 需要海量的质检工作,而很多组织根本不肯投这份钱。APEX 是一个 RAG benchmark,做法是把文件交给 agent,然后就文件内容提问。但在有些样例里,文件里实际写的东西和评分标准里期望的答案对不上。于是,agent 要是老老实实照着它看到的 ground truth(标准答案依据)去做,反而会拿到负分。
[10:05] Nick Heiner
And a lot of the data in Apex is seemingly synthetically generated because it's full of obvious placeholder values or dates or places that don't exist. And so as a result, the model is more likely to develop eval awareness where it realizes that it's being tested which undermines the entire exercise. It also just takes you out of distribution from actual real world data to something that is obviously fake. So that's an overview of some of the key antiatterns that happen during benchmark creation. But benchmaxing is a two-way process and there are all sorts of fun things that labs can do to benchmax and that's what we're going to talk about next. So the the core value that we're all trying to get towards as human eval right AI exists to serve humans and so just having humans look at the responses and make ratings like that's what we care about. The problem is that human eval is very expensive. And so a lot of what benchmarks are doing is trying to get around that and you are trying to distill human preference into something more scalable and you're hoping you do that distillation in a way that's still sufficiently faithful to what human eval wants. But what this means is that inevitably there is a point where you can keep hill climbing on a benchmark and the human eval stays flat.
而且 APEX 里很多数据看上去是合成生成的,因为里面全是明显的占位符数值、根本不存在的日期或地名。这带来的后果是:模型更容易产生 eval awareness(评测意识)——它意识到自己正在被测试,而这就把整场测试的意义给掏空了。同时它也让你彻底偏离了真实世界数据的分布,跑到一堆一眼假的东西上去了。以上就是 benchmark 制作环节里几个关键反模式的概览。但 benchmaxxing 是一个双向的过程——实验室那边也有各式各样好玩的手段可以用来刷榜,这就是我们接下来要聊的。我们所有人真正想逼近的核心价值,其实是 human eval(人类评测)——AI 存在的意义是服务人类,所以让人类直接去看模型的回答、去打分,那才是我们真正在意的东西。问题是 human eval 非常贵。所以 benchmark 做的很多事,本质上都是在绕开这个成本:你想把人类偏好蒸馏成某种更可规模化的东西,同时期望这次蒸馏足够忠实于 human eval 想要的东西。但这也就意味着,必然会出现一个临界点——在那之后,你在 benchmark 上继续爬坡,human eval 的曲线却已经躺平了。
[11:27] Nick Heiner
And you can actually take it even further if you want where you keep hill climbing on a benchmark even as the human eval goes down. But if for whatever reason you think this is necessary for marketing or we have sort of organizational politics or incentives that are demanding this that's how it can end up happening. In this instance the prompt is what time is it? And the response is absolutely deranged. No human eval is ever going to choose this but El Marina puts it at the top of the leaderboard. So again, you have this divergence and if you're trying to benchmax, you just cannot care about that. Another thing you can do that I've heard stories of is you can actually hire a crowdsource army to vote for you in Elmarina since Elmarina basically does no filtering of their workforce. And you might say, well, we anonym, you know, Elmarina anonymizes. So how are they going to know who to vote for? That's actually quite simple. You have your model include a watermark that tells the crowd who to vote for. There's also all sorts of things you can do with running your evals in conditions that are like not fully representative of the applesto apples comparison you're trying to make and then not always being super transparent about those conditions in such a way that undermines the validity that the community is trying to interpret because they don't have that contextualizing information.
如果你愿意,你甚至可以更进一步:benchmark 分数继续往上爬,而 human eval 是往下掉的。但只要出于某种理由——你觉得营销上必须这么干,或者组织内部的政治和激励在逼着你这么干——事情就会这样发展。看这个例子,prompt 是「现在几点了?」,而模型的回答简直精神错乱。没有任何一个 human eval 会选这条回答,但 LMArena 把它排到了 leaderboard(排行榜)第一。所以你又看到了这种背离;而如果你的目标就是 benchmaxxing,你就只能对这种背离视而不见。还有一招我听说过:你可以直接雇一支众包大军去 LMArena 给你投票,因为 LMArena 对参与投票的人群基本不做任何筛选。你可能会说,LMArena 不是做了匿名化吗,他们怎么知道该投给谁?这其实很简单——让你的模型在输出里带上一个水印,告诉这群人该投哪一个。此外还有各种花样:你可以在并非完全对等的条件下跑自己的 eval,让这场比较其实算不上 apples-to-apples(同口径对比),然后在披露这些条件时不那么透明。这样一来,社区在解读结果时就失去了那些用于还原上下文的信息,整个结论的有效性也就被架空了。
[12:47] Nick Heiner
This was a paper um again about Elmarina and talking about how some of the dynamics of how it's run lead to models overfitting on Elmarina. Um in this instance, the specific chart we're seeing is that Meta tested 27 models without disclosing that it was doing so. Um which you know distorts the results. So how are we going to end benchmaxing? We need to hold the benchmark industry and the labs to a higher standard. The first thing we need to do when making a good benchmark is start with great human experts. And those experts inform everything that is downstream from what types of tasks are we going to have the agent do? How is success measured? What are the input files that agents are given? What are the tools that they're given? But we also do need that product sense. So imagine you're making a medical benchmark. It's not enough to have doctors who can answer specific medical questions because if you're trying to test how ready are we for agents to be deployed into hospitals. You also need someone with the business sense to know what's the regulatory environment, what's the legal requirements because that is going to impact what types of tasks you're trying to have the AI solve. You need high fidelity input data which is best done by going out and getting it from the real world, having actual people create this data. Synthetic approaches are possible, but it is very very hard to do it reliably.
这是一篇论文,同样是讲 LMArena 的,讨论它的运作机制里有哪些动力学会导致模型在 LMArena 上 overfitting(过拟合)。我们现在看到的这张图具体说的是:Meta 测了 27 个模型,却没有披露自己在这么做——这显然会扭曲结果。那么,我们要怎么终结 benchmaxxing?我们必须用更高的标准去要求 benchmark 行业和各家实验室。做一个好 benchmark,第一件要做的事是从一流的人类专家开始。这些专家会决定下游的一切:我们要让 agent 做哪类任务?成功怎么衡量?给 agent 的输入文件是什么?给它的工具是什么?但同时,我们也确实需要产品感。想象你要做一个医疗 benchmark,光有能回答具体医学问题的医生是不够的——因为如果你真正想测的是「我们离把 agent 部署进医院还有多远」,你还需要一个懂业务的人,知道监管环境是什么样、法律要求有哪些,因为这些会直接决定你要让 AI 去解决哪些类型的任务。你还需要高保真的输入数据,而拿到它的最好办法就是去真实世界里取材,让真实的人来创造这些数据。合成的路子不是不行,但要做到可靠,非常非常难。
[14:14] Nick Heiner
The tools need to actually work. A lot of benchmarks have tools that are buggy in various ways. And unless you're intentionally making a benchmark about buggy tools, this just introduces noise. You need verifiers that are fully aligned with the prompts. And this is a two-way alignment. So the verifiers need to be verifying everything the prompt asks for. And everything the prompt asks for needs to be covered by the verifiers. And if you get either side of those two misaligned, then it's going to be unfair to models and you're introducing random noise. You need to thoroughly QC everything and you need to have a private hold out set so you don't get contaminated. And if you do all this right, then you'll avoid what often happens with benchmarks, which is when labs get to like 80% and say, "Okay, this is saturated." And I used to think that saturation was just them saying again we don't think training on this further is going to increase real world value. And it often does mean that but it can mean that because the lab is saying we realize 20% of these tasks are broken. But the problem is that as you're hill climbing you don't know what 20% are broken until you solve all the others. And so as a result you have a lot of noise. And if that 20% of broken tasks is randomly but in a biased way assigning the rewards, it's going to really distort the model relative ranking you're trying to get.
工具必须真的能用。很多 benchmark 里的工具都有各种各样的 bug。除非你就是故意要做一个「考验模型怎么应付坏工具」的 benchmark,否则这只会引入噪声。你需要 verifier 和 prompt 完全对齐,而且这是双向的对齐:verifier 要把 prompt 要求的每一件事都验到,prompt 要求的每一件事也都得被 verifier 覆盖到。这两个方向只要有一边没对上,对模型就是不公平的,你就是在往里灌随机噪声。你得把所有东西都彻底 QC 一遍,还得留一份私有的 hold-out set,这样才不会被污染。如果这些你都做对了,就能避开 benchmark 上经常发生的那种事——各家实验室做到 80% 左右就说:「行了,这个已经饱和了。」我以前一直以为「饱和」的意思无非是他们觉得再往这上面训也不会带来更多真实世界的价值。它确实经常是这个意思,但它也可能意味着:实验室发现这些任务里有 20% 本身就是坏的。问题在于,你在爬坡的过程中,根本不知道是哪 20% 坏了——非得把其余的都解完才看得出来。结果就是一大堆噪声。而如果那 20% 的坏任务是在以一种随机但有偏的方式分配奖励,那它会严重扭曲你想测出来的模型相对排名。
[15:41] Nick Heiner
So at Serge, we created a benchmark called Hemingway bench to measure writing. There have been a number of writing benchmarks that use various mechanical means to try to assess writing quality, but we believe that writing is just too rich and deep and nuanced and frankly human of an activity to measure with mechanical benchmarks and LM as a judge doesn't really work either because LLMs don't have good taste in writing. Again, this is sort of the you can't expand the frontier from within the frontier situation. So what we've done is we've just created a workforce of thousands of professional writers in various domains, technical writers, poets, journalists, editors, and we just have them do blind model comparisons and then we create this leaderboard and it is quite expensive, right? Human eval is very expensive. Getting the time of these professionals is quite expensive. But again, our goal is to maximize quality, not to minimize costs. So in conclusion, benchmaxing is the exploitation of benchmark misalignments between human preference, but we can do better and we can hold the industry to a higher standard. Both the people making the benchmarks like myself and the people who are reporting on the benchmarks. And if you'd like to be a part of that, of course, obligatory pitch at Serge, we're hiring for basically all aspects of that. Uh and if you'd like more spicy takes from me, uh please follow my substack. Thank you very much.
所以在 Surge,我们做了一个叫 Hemingway Bench 的 benchmark 来衡量写作。此前已经有不少写作类 benchmark,用各种机械化的手段去评估写作质量,但我们认为,写作这件事太丰富、太深、太微妙,说白了太「人」了,没法用机械化的 benchmark 来衡量;而 LLM-as-a-judge 也不太行,因为 LLM 在写作上没什么品味。这又回到那个「你没法从边界之内去拓展边界」的处境。所以我们的做法是,组建了一支上千人的专业写作者队伍,覆盖各个领域——技术写作者、诗人、记者、编辑——让他们做盲测的模型对比,然后我们据此做出这个榜单。这当然很贵,对吧?人工 eval 非常贵,占用这些专业人士的时间也非常贵。但还是那句话,我们的目标是把质量做到最高,而不是把成本压到最低。所以,总结一下:benchmaxxing 就是在利用 benchmark 与人类偏好之间的错位。但我们可以做得更好,我们可以用更高的标准要求这个行业——既要求做 benchmark 的人(比如我自己),也要求那些拿 benchmark 说事、做报道的人。如果你想参与其中,那当然少不了一句例行安利:Surge 正在招人,基本上上面这些方向全都招。另外,如果你想看我更多「重口味」的观点,欢迎关注我的 Substack。非常感谢大家。