ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.129 · 全文

What's Next After RLHF? — Diogo Almeida, TypeSafe AI

频道: AI Engineer
视频: https://www.youtube.com/watch?v=cJ0EOzey--o
原文语言: en
统计: 共 13 轮 · Diogo Almeida 12 · — 1


[0:12] Diogo Almeida

Excellent. I will say that um I might speed run through this. Feel free if you don't disag- agree with something to yell out. It's way more fun for me if things get interactive. Um otherwise, I will go through this. Uh first, can I have like a vague show of hands of who knows what RLHF is? Oh, excellent. I might be able to skip through that part quickly and get into the interactive stuff. So, my name's Tiago Almeida. I'm talking about what's next after RLHF. More accurately, I think this should be called what's next after the chat GPT era that I think we're all in. And my hint for you guys is it is not the Claude code era. I will justify this later on, but I actually believe them to be part of the same era. Why should you listen to me? I was co-authored to what what is basically OpenAI's greatest hits, at least published hits. Co-authored to GPT-4, chat GPT, RLHF {slash} instruct GPT. Um the team I was part of basically invented post-training as a concept. So, um very qualified on a lot of this stuff. But what makes me somewhat unique here is that I'm one of the few people at OpenAI who actually hates on chat GPT. Uh thank you. Uh I don't hate chat GPT as a product, to be clear. I think chat GPT is a world-changing product that will probably stay with us for the rest of time unless something better comes up. But I also acknowledge its limitations and I I I think a lot of what's happened in the state of the field can be traced back to minor decisions we made in making the algorithms behind chat GPT.

太好了。先说一句,我可能会讲得比较快。哪里你不同意,随时喊出来——对我来说互动起来有意思多了。不然我就自己一路讲完。第一个问题,大概举个手,知道 RLHF 是什么的?哦,很好,那这部分我可以快速跳过去,直接进到好玩的部分。我叫 Diogo Almeida,今天讲的是「RLHF 之后是什么」。更准确地说,我觉得这个题目应该叫「我们现在所处的 ChatGPT 时代之后是什么」。先给你们一个提示:不是 Claude Code 时代。后面我会论证,但我其实认为它俩属于同一个时代。为什么该听我讲?OpenAI 那批「金曲精选」——至少是公开发表的那些——我基本都是共同作者:GPT-4、ChatGPT、RLHF / InstructGPT。我所在的那个团队,基本上算是发明了 post-training(后训练)这个概念。所以这些东西我还是挺有资格讲的。但我在这儿有点特别的地方是:我是 OpenAI 里少数几个会吐槽 ChatGPT 的人。谢谢。说清楚,我不是讨厌 ChatGPT 这个产品。我觉得 ChatGPT 是个改变世界的产品,除非出现更好的东西,否则它大概会一直陪着我们。但我同时也承认它的局限。而且我认为,这个领域今天的很多状况,都能追溯回我们当年做 ChatGPT 背后那套算法时做的一些很小的决定。


[1:49] Diogo Almeida

Um I feel like the question that's relevant to everyone in AI right now is what's actually going on. Um there's a lot of like differing opinions, and I think it's really useful to like map out the spectrum and figure out how can smart people have like such different opinions. There's cult one. Um, AI is not just going well, it's going insanely well. Every single benchmark we surpass human level, and as far as we can measure, we are continuously surpassing human performance. You know, like basically every new benchmark, and it's only getting faster and accelerating. You have uh, you know, every Can I see my mouse? Excellent. Basically every like NLP benchmark is getting crushed, and not only that, allegedly the time that LLMs can operate autonomously is growing exponentially. On the other hand, you have AI is not just going poorly, it's going like insanely poorly. AI is a bubble, it's basically generating no value, it's just circular financing deals, etc., etc. And, you know, if AI is so great, why is why is everything just like a chat app right now? Or like a cloud go thing? Um, and a a lot of the people have actually kind of given up on what was the old guard's terminology of a transformative AI revolution. People aren't really talking about that anymore. They're talking about it being like massively valuable like B2B SaaS.

我觉得现在 AI 圈里每个人都关心的问题是:到底在发生什么?大家的看法差得非常远,我觉得把这个光谱摊开来画一画、搞清楚为什么聪明人能有这么不一样的判断,特别有用。第一个教派:AI 不只是发展得好,是好得离谱。每一个 benchmark 我们都超过人类水平,而且就我们能测的范围看,我们在持续超越人类表现。基本上每出一个新 benchmark 都这样,而且速度还在加快、还在加速。你看——我的鼠标在哪儿?好,找到了——基本上每一个 NLP benchmark 都被碾压了。不只如此,据说 LLM 能自主运行的时长还在指数级增长。另一边呢,是 AI 不只是发展得差,是差得离谱。AI 是个泡沫,基本没创造什么价值,全是左手倒右手的循环融资,等等等等。而且你想,如果 AI 这么牛,为什么现在所有东西看着都只是个聊天 App?或者一个 Claude Code 那样的东西?还有,很多人其实已经放弃了老一辈那套「变革性 AI 革命」的说法了。现在没人真的在谈那个了,大家谈的是它会是一门「巨值钱的 B2B SaaS」。


[3:12] Diogo Almeida

So, the only thing that everyone agrees on is like there's just a these extreme points of view and like nothing in between. And everyone basically thinks AI is insane, but like for different reasons. And what I would want to talk about is what is the sane view of AI? Let's take all the evidence of like cult one, it's going super well. Take all the evidence of cult two, it's going super poorly. Like, uh, you know, map them out and try to explain what what what explains that divide. Like, what is the simplest possible explanation of why some things are too good to be true, and some things are not just bad, they are so bad that we would still employ human workers to do like, you know, like kind of like dumb tasks. Um, no offense to any of them. A lot of these tasks on the right seem way, way, way easier than the stuff on the left. Like, how can we be solving like, you know, unsolved math problems, but still customer service requires like humans in the loop in order to actually like make decisions? This I think is like a kind of like a wild state of affairs. And in my opinion, anyone who works adjacent to AI should have an answer to this because this is like the evidence in the field right now.

所以唯一大家都同意的一点是:只有这两个极端,中间什么都没有。所有人都觉得 AI 疯了,只不过疯的理由不一样。我想聊的是:那个清醒的 AI 观到底是什么?把第一个教派的证据全拿过来——发展得超好;把第二个教派的证据也全拿过来——发展得超烂。把它们摆在一起,试着解释这个分裂到底怎么回事。什么是最简单的解释,能说明为什么有些事好得不像真的,而有些事不只是差,是差到我们宁可还雇人类去干那些挺蠢的活儿。没有冒犯任何人的意思。右边这些活儿看着比左边那些容易太多太多了。我们怎么可能一边在解还没解开的数学难题,一边客服却还得要人类在回路里做决定?我觉得这个局面挺离谱的。而在我看来,任何跟 AI 沾边工作的人,都该对这件事有个答案,因为这就是这个领域眼下摆在桌上的证据。


[4:20] Diogo Almeida

Um, I would normally pause and ask people if they want to like yell out their thoughts in this, but uh, that I don't think we have time for that and I've been told to not take Q&A until after. Um, but I'll just give you my answer to this, which is, in my opinion, the simplest explanation. All the stuff on the left is not just a task that happens to have a human in the loop. In the left, the task The goal of it is to please the human in the loop. These tasks are intrinsically human in the loop tasks. The like Claude code's job is not to just make code work. Um, the the the the way it converses would be totally different. The goal is to please the human in it. And on the other side, all of these tasks that seem way more basic, the goal is to not have remove the human loop. Ideally, it would be running in the background in a server that you never even look at and ideally it eventually becomes like legacy software that you don't really worry about. So, and this is the divide between assistance and automation. Um, lesson one for my talk is that today's AI, everything inherited from our LHF, is incredible at the human in the loop stuff, but not for automation tasks. Tasks. This is a longer side, but the lesson basically every business has learned is do not use AI for decisions with stakes to your business. Um, a common pattern is make sure that all of the costs are to the user and not to your business. So, um, it's oh, totally okay to throw the user at infinite docs in customer service, but it is not okay to make it make expensive decisions.

平常讲到这儿我会停下来,让大家喊出自己的想法,但我们时间不够,而且主办方让我把 Q&A 留到最后。那我直接给我的答案,我觉得这是最简单的解释:左边那些事,不是「碰巧有个人在回路里」的任务。左边那些任务,它的目标本身就是取悦回路里的那个人。这些任务本质上就是 human in the loop(人在回路)的任务。Claude Code 的活儿不只是把代码跑通——如果只是那样,它说话的方式会完全不同。它的目标是取悦回路里那个人。而另一边,那些看着基础得多的任务,目标恰恰是把人从回路里拿掉。理想状态是它在某台服务器上后台跑着,你看都不用看一眼;最理想的是它最后变成你根本不用操心的 legacy 软件。这就是 assistance(当助手)和 automation(自动化)之间的分野。我这场的第一课:今天的 AI——所有从 RLHF 继承下来的东西——在人在回路的事情上强得不得了,但对自动化任务不行。这一课还能再展开,但基本上每家公司学到的教训是:不要把 AI 用在对你业务有真实赌注的决策上。一个常见的做法是,确保所有代价都落在用户身上,而不是落在你的业务上。所以,在客服里把用户扔进无穷无尽的文档里,完全 OK;但让它去做那种代价昂贵的决策,就不 OK。


[5:49] Diogo Almeida

Horrible pattern, but that is the state of AI right now. Uh I can I can blitz through the what is RLHF part cuz you all seem to know what it what it is. Um it's the algorithm behind not just ChatGPT, but basically every LLM today. As far as I can tell by usage, 100% roughly of LLMs are trained with RLHF. And we have this we as in we the OpenAI team had this great blog post on how it worked. Um I will not get into that because you all know it, and this is super boring. Um the summary of this is it is just collect human preferences, optimize for human preferences. Um and if you want to see like an annotated version of this, you can see which parts are collecting human preferences, which ones are optimizing for them. And [snorts] this, I think, provides a really clear answer to everyone in the field asking, "Why do all LLMs require a human in the loop?" The And the simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference. It is not to run software autonomously. It's kind of super obvious. Thank you, my man at the back. The Yeah. I I I love that you're laughing at this. Um and because of that, overpromising is a feature. This is by design. This is an old meta study. Um and the the numbers probably have changed, but by construction, every RLHF model will always have a big difference between human preference and results, even if the results are good, because the main objective you're optimizing for is for human preference.

这个模式很糟糕,但这就是 AI 现在的状态。「什么是 RLHF」这部分我可以飞快过掉,因为你们好像都知道。它不只是 ChatGPT 背后的算法,基本上是今天每一个 LLM 背后的算法。就我看到的使用情况,大概 100% 的 LLM 都是用 RLHF 训出来的。我们——我是说我们 OpenAI 团队——写过一篇特别好的博客讲它怎么运作。这块我就不展开了,你们都知道,而且讲起来无聊透了。总结就是:收集 human preference(人类偏好),然后照着 human preference 优化。你要是想看标注版,能清楚看出哪部分在收集人类偏好、哪部分在为它做优化。而这,我觉得,给这个领域所有人都在问的那个问题一个非常清楚的答案:「为什么所有 LLM 都需要一个人在回路里?」简单答案是:我们真的就是把人放进回路里了。这个回路的目标就是为 human preference 做优化,它的目标从来不是让软件自主运行。这不是明摆着的吗。谢谢后排那位兄弟。我特别喜欢你笑这个。也正因为这样,over-promising(过度承诺)是个 feature,是设计出来的。这是 Meta 一份挺老的研究了,具体数字大概已经变了,但从构造上讲,每一个 RLHF 模型的「人类偏好」和「实际结果」之间永远会有一个大落差——哪怕结果本身是好的——因为你主要优化的目标就是人类偏好。


[7:24] Diogo Almeida

This is just like natural to how LLMs work. Um I love this tweet of um uh sending ChatGPT an audio file of fart sound effects and asking like what What do you think of the music I made? Here's a straight honest reaction. It's a very eerie vibe atmosphere piece. And this is just how RLHF works. If it doesn't know, it will err on the side of doing what it thinks is best for human preference. And this makes total sense if you are a user in the loop because like the end game for all RLHF models is optimizing for engagement. But what you really want if you want automation is for it to just like not give a about the humans and just do the task correctly in a calibrated way. Um Lesson number two is that today's AI was designed for assistance through optimizing for human preference. This is like it's like in the name. This is not like a controversial take. And the consequences are maybe more controversial, but it's like very obvious if you think about what we really are optimizing for, which is no matter how wrong the models are, they will look right because of the asymmetry within the reward model in RLHF. Um and this is where a lot of like the dilemma in the field stems from because people really want automation to happen.

这就是 LLM 的天性。我特别喜欢这条推:有人给 ChatGPT 发了一个放屁音效的音频文件,问它「你觉得我做的这段音乐怎么样?给我一个直接、诚实的反应」。它回答说:这是一段氛围极其诡谲的意境作品。这就是 RLHF 的运作方式。它要是不知道,它就会往「它认为对人类偏好最有利」的那一边偏。如果你是回路里的那个用户,这完全说得通,因为所有 RLHF 模型的终局就是优化 engagement(互动性、粘性)。但如果你要的是自动化,你真正想要的是它压根别管人类怎么想,就把任务做对、做得 calibrated(校准得当)。第二课:今天的 AI 是为 assistance 设计的,方式就是优化 human preference。这名字里就写着呢,这不是什么有争议的说法。它的后果可能就比较有争议了——不过你只要想想我们真正在优化的是什么,其实非常明显:不管模型错得多离谱,它看起来都会是对的,因为 RLHF 里的 reward model(奖励模型)存在一种不对称。这个领域里很多两难就是从这儿来的,因为大家是真的很想要自动化。


[8:47] Diogo Almeida

Cool. So, back to the original question. I'm over halfway done with the talk and I haven't even answered it. I was just talking about what's RLHF. But this was a framing to talk about what RLHF is to talk about what's next. And I would say the real question is what's next after AI's assistance era, which I think that we are like very firmly in right now. And back to the original clue of why it's not Claude code. It's actually a super fun nuanced discussion, but it's not Claude code because Claude code is still part of that assistance era. Claude code is still RLHF and it'll it would look very very different if it was purely This is a little advanced, but if it was purely RLVR, it would look very very different. And this is why you get like this dilemma with models where sometimes it gets really good at agentic stuff, but it stops following what you actually want. This This like the trade-off in optimization space that keeps dancing, but both of these trade-offs in optimization space do not add to the automation component. And like that leads to what I think the the logical answer of what's next after assistance is real automation. Um to talk about a little bit about the automation and how that would work, I want to talk about software. Um maybe this is a little bit philosophical for you guys, but I think it's when it clicks and hopefully it clicks if I do a good job. It it I I hopefully it'll be like really clear, which is I'm a lover of software. I assume everyone here loves software. Software is like super valuable. See all the SaaS. And kind of like the craziest part of software in my opinion is that all of the SaaS basically has not changed since 2019.

好。回到最开始那个问题。我这场已经讲过半了还没回答它,一直在讲 RLHF 是什么。但这是个铺垫——先讲清楚 RLHF 是什么,才好讲接下来是什么。我会说,真正的问题是:AI 的 assistance 时代之后是什么?而我觉得我们现在非常牢固地就在这个时代里。回到最开始那个「为什么不是 Claude Code」的伏笔。这其实是个特别有意思、也很微妙的讨论,但它之所以不是 Claude Code,是因为 Claude Code 仍然属于那个 assistance 时代。Claude Code 仍然是 RLHF。如果它是纯 RLVR 的——这个稍微进阶一点——它会长得非常非常不一样。这也是为什么你会碰上那个两难:模型有时候 agentic(智能体式)的活儿干得特别好,但它就不听你实际想要的了。这是优化空间里来回摇摆的权衡,但这两头的权衡,哪一头都没有往「自动化」这个维度上加分。这就引向了我认为合乎逻辑的那个答案:assistance 之后,是真正的 automation。要聊自动化、聊它会怎么运作,我想先聊聊软件。这段对你们来说可能有点哲学,但我觉得这是会「啪」一下想通的那种东西,希望我讲得够好、你们能想通。我是个软件爱好者,我假设在座各位也都爱软件。软件超级值钱,看看那一堆 SaaS 就知道了。而在我看来,软件最疯的一点是:所有这些 SaaS,从 2019 年到现在基本没变过。


[10:23] Diogo Almeida

Like SaaS is not really changed in the LLM era, except sometimes a chatbot is like latched on, which is like kind of insane if you think about like the progress made in AI, but is actually very predictable when you think that AI is assistance native, right? Like AI is made for assistance. What can you do in SaaS? Just provide an assistant on the side. And this is not what early AI pioneers used to think would happen. Like when you see like the early wording in opening eyes uh charter, it's about like doing like tons of work, not about like making profit or anything like that. And we used to think that software would get a lot smarter, not just cheaper to write, which is kind of the direction we're going down right now. And I actually really like this phrasing from Garry Tan. Um I think he means this as a compliment to what's going on right now. We're entering the golden age of just-in-time software, but I actually think that this is like a like a double-edged sword. Like I don't just want just-in-time software, which is cool. I I love cloud code, to be clear, just like I love ChatGPT. I would keep using it. But like what I want is smarter software. Why can't like B2B Why can't software just be more expressive? Like why are the like the building blocks of software actually still the same? And um I think this is a question that the whole AI industry should ask itself. And basically every time you're thinking about we want to do automation, it is not about like, you know, an amalgamation of like automating a person's work. It's about like, "Hey, there's this extremely rote work. It's so simple that we can like communicate to someone else that this thing should be done." And ideally it's like it it's so basic that it could be done repeatedly for basically free.

SaaS 在 LLM 时代基本没变,顶多是旁边挂了个 chatbot。你要是想想 AI 这几年的进展,这挺离谱的;但你要是意识到 AI 是 assistance-native(生来就是做助手的),这又完全可以预料。AI 就是为助手场景造的。那你在 SaaS 里能干什么?在边上挂一个助手呗。而这不是早期那批 AI 先驱以为会发生的事。你去看 OpenAI 章程(charter)里最早的措辞,讲的是「干掉大量的工作」,不是讲赚钱之类的。我们当年以为软件会变得聪明得多,而不只是写起来更便宜——而后者恰恰是我们现在正在走的方向。我其实挺喜欢 Garry Tan 这个说法。我觉得他说这话是在夸现在这个局面:我们正在进入 just-in-time software(即用即造的软件)的黄金时代。但我觉得这是把双刃剑。我要的不只是即用即造的软件——虽然它确实很酷,说清楚,我是真的很爱 Claude Code,就像我爱 ChatGPT 一样,我会一直用下去。但我想要的是更聪明的软件。为什么软件不能更有表达力?为什么软件的那些基本积木到今天还是原来那几块?我觉得这是整个 AI 行业都该问问自己的问题。而且基本上每次你在想「我们要做自动化」的时候,那件事不该是把某个人的工作胡乱拼一拼自动化掉,而该是:「嘿,这儿有一件极其机械重复的活儿,它简单到我能直接讲清楚让别人去做。」而且理想情况下,它基础到可以被反复执行、几乎零成本。


[12:06] Diogo Almeida

Um or it could be done by computers. And that's really not happening right now. What we're doing is we're just automating the writing of the software. But then it its expressibility is the same. And that's That to me is like tragic in the state of the world. Um Cool. Oh, lesson three. Um This is something that I believe strongly in. I believe that like eventually the field will write the I wouldn't say Arlatech is a wrong, but it was like a weird detour and one that we didn't expect. Tomorrow's AI, I believe, will be for automation. And we will eventually have a world with smarter software. Like there will start to be actual work that is automated, which I, you know, right now it's a rounding error despite LLM's intelligence. And that is what we are working on at TypeSafe. We are still kind of stealthy. Like I'm willing to give these talks, but these are like some of the early ones. Um Our core question is what if the AI stack was redesigned for reliability and automation? Like how would that all change? What what would you do? And actually there's a lot It's a It's a very interesting fork in the road for what's go you know, like from basically every LLM that's built today. And I think it's one of the most satisfying things I've worked on, and I've worked on some pretty cool stuff. We are releasing soon. So, um if you want to work with us or you want to like you know, be the first one of the first to build smart software, please sign up on either our mating mailing list or careers page. And I am trying to start a Twitter. So, follow me and I will post really spicy things. I actually will post something later today that I guarantee will be very spicy. Uh the hint is that the original scaling laws were incorrect.

或者说,它可以交给计算机来做。而这件事现在根本没在发生。我们现在做的只是把「写软件」这件事自动化了,但软件本身的表达力还是原来那样。这在我看来是这个世界现状里挺悲哀的一点。好,第三课。这是我非常坚信的一点。我相信这个领域最终会——我不会说 RLHF 是错的,但它是一段奇怪的弯路,一段我们没预料到的弯路。我相信,明天的 AI 会是为 automation 服务的,我们最终会拥有一个软件更聪明的世界。会开始出现真正被自动化掉的工作——而现在,尽管 LLM 这么聪明,这部分几乎可以忽略不计。这也正是我们在 TypeSafe 做的事。我们现在还处在半隐身状态,我愿意出来讲这些,但这属于最早的几场之一。我们的核心问题是:如果整个 AI 技术栈是围绕可靠性和自动化重新设计的,会怎么样?那一切会怎么变?你会怎么做?其实里面能岔出去的东西很多,相对于今天所有 LLM 的建法,这是一个非常有意思的岔路口。这也是我做过的最带劲的事情之一,而我做过的东西还挺酷的。我们很快就会发布。所以,如果你想跟我们一起干,或者你想成为最早一批做「聪明软件」的人,欢迎去我们的邮件列表或者招聘页登记一下。另外我在试着搞 Twitter,关注我,我会发一些很猛的东西。今天晚点我就会发一条,我保证特别猛——提示一下:最初那版 scaling laws 是错的。


[13:53] Diogo Almeida

Cool. Um that uh that's it for my prepared stuff. I would love Do I have time for for people yelling out questions? I would love questions, feedback, disagreements, strong stuff. I can repeat the question. You don't have to worry about the mic. Hell yeah. Uh cool. Uh the question was roughly what if you trained like a classifier head with pre-training as well? Uh roughly uh like Yoshua Bengio is suggesting. Um I will say that that's complicated. And I actually think I don't have the time to answer that particular question. I will give like my simplified view on this. And it the answer is I actually don't think that pre-training is the problem. I think pre-training is uh phenomenal. Like the fact that we compress the knowledge of the internet into like this core of intelligence that then can be utilized is incredible. And the pre-trained models are incredibly intelligent. Uh and I believe that the problem is like how we unearth it. And hallucination [clears throat] to me is intrinsic to um optimizing for human preference. Like there's an asymmetry in the reward model kind of like a GANs have. Oh, I really should not get This is a very advanced topic. But there's an asymmetry in the reward model like what GANs have that allow for um that encourage the models to drop modes and be confident because it's very easy to see when the model is not confident and to punish that from a reward model perspective. It's very complicated, but uh I'm happy to chat afterwards if you want to jam.

好,我准备的内容就到这儿。我很想——我还有时间让大家喊问题吗?我特别欢迎问题、反馈、反对意见,越猛越好。我可以复述问题,你不用管话筒。太好了。这个问题大概是:如果在 pre-training(预训练)阶段就顺带训一个 classifier head(分类头)会怎么样?大致就是 Yoshua Bengio 提的那个思路。我得说这事很复杂,而且我觉得我没时间回答这个具体问题。我给一个简化版的看法:我其实不认为 pre-training 是问题所在。我觉得 pre-training 好得惊人。我们能把整个互联网的知识压缩成这么一个可以被调用的智能内核,这件事本身就不可思议。预训练出来的模型智能得吓人。我认为问题在于我们怎么把它挖出来。而 hallucination(幻觉)在我看来,是「为 human preference 做优化」这件事天生带来的。reward model 里存在一种不对称,有点像 GANs 里的那种。哦,我真不该往下讲——这话题太进阶了。但 reward model 里确实有一种类似 GANs 的不对称,它会让模型 drop modes(丢掉某些模式)、并且表现得很自信,因为从 reward model 的角度看,「模型不自信」是非常容易被看出来、也非常容易被惩罚的。这事挺复杂的,不过你要是想深聊,散场后我们随时聊。


[15:25] Diogo Almeida

Cool. Oops. Um I have other slides from other talks as well that I could go into more about that. I have a minute left. Hell yeah. Say it again. It is definitely not RLVR. So, it is a new thing. Every single optimization stack I will actually go into an old presentation that I have because I think this is super important. Um in terms of like to me what the like Sutton's bitter lesson is that algorithms matter more than compute. This is true in games, but not true in reality. I actually think that the full stack is that data matters more than compute and doing the right task matters way more than data. And basically every single branch of LLM post-training if you want to call it has its own North Star of what it's optimizing for. So, RLHF is optimizing for human preference. RLVR is optimizing for like log error rates of pure correctness, but we are doing a third thing that is optimized for calibrated decision-making and like basically mainlining the intelligence of pre-trained models into like being actually useful for software, which I think is like quite different. Uh could you say that again? Uh they're asking if the the reward is injected through the whole process. I will actually say that even the shape of the API is different because the shape of the API for RLHF is different from RLVR, which is different from what we are doing. So, we are like thinking about it from scratch just like no one thought about instruction following before we made instruction following happen. Um usually when there's a big branch in new ways to post-train, like it it it just looks like totally alien, and then in hindsight becomes super obvious.

好。哎呀。嗯,我还有别的演讲里的幻灯片,可以就这块再多讲一点。我还剩一分钟。太好了。你再说一遍。那绝对不是 RLVR,所以这是个新东西。每一套优化栈——我干脆翻回我以前的一个演示,因为我觉得这一点特别重要。嗯,在我看来,Sutton 的 bitter lesson 讲的是算法比算力更重要——这在游戏里成立,但在现实里不成立。我其实认为,完整的那套栈是:数据比算力更重要,而做对任务又远比数据更重要。基本上,LLM post-training(如果你要这么叫的话)的每一个分支,都有它自己的北极星指标。RLHF 优化的是 human preference,RLVR 优化的是纯正确性那种 log error rate,而我们在做的是第三件事,优化的是 calibrated decision-making,基本上就是把预训练模型里的智能直接灌进去,让它对软件真正有用——我觉得这挺不一样的。呃,你能再说一遍吗?呃,他们问的是 reward 是不是贯穿整个流程注入的。我甚至想说,连 API 的形态都不一样:RLHF 的 API 形态跟 RLVR 不同,跟我们在做的又不同。所以我们是从零开始重新想这件事,就像在我们把 instruction following 做出来之前,也没人想过 instruction following 一样。嗯,通常一旦 post-train 出现一条全新的大分叉,它一开始看上去完全像外星玩意儿,事后回头看却特别显然。


[17:22] Diogo Almeida

Cool. I believe I'm overtime cuz this red thing is is beeping, but please find me afterwards. I love questions. I love the interactivity. Um and uh follow me on Twitter for spicy stuff. Heck, yeah. Oh, oh yeah, it's over here. Complete skeptic. Um it's it's on brand for me. Cool. Heck, yeah. Thank you.

好。我估计已经超时了,因为那个红灯在滴滴响,但散场之后一定来找我聊。我特别喜欢被提问,喜欢这种互动。嗯,还有,想看点带劲的内容就在 Twitter 上关注我。好嘞。哦,哦对,在这儿呢。彻头彻尾的怀疑论者——嗯,这挺符合我人设的。好。太好了。谢谢大家。


[18:01]

[music]

[music]