8 Predictions for the Era of Continual Learning
频道: Dwarkesh Podcast
视频: https://www.dwarkesh.com/p/era-of-continual-learning
原文语言: en
统计: 共 10 轮 · Dwarkesh Patel 10
[0:00] Dwarkesh Patel
So I've explained elsewhere why I think actual continual learning is needed.I don't think you can have AIs that perform whole jobs as competently as humansif they are forced to just write marked-on files from session to session.Just to give an illustrative example,imagine if this is the way that students learn to play the saxophone.So you have one student, he's never played the saxophone before,he goes into the music hall, he tries to play it, of course this is his first time,so he fails.And he writes down a bunch of notes about what went wrong.And there's a next student who's waiting outside the music hall,he comes in, he reads all these notes,he's also never played, so of course he messes up,and he continues to add on to these notes.And you have an infinity of students who are outside the music hallwho keep writing notes to the next person.I don't think there's any sequence of text they could write to each otherthat would allow the subsequent student to just nail the saxophone from the first try.At some point, you actually have to accumulate the relevant experience into your brain.I think the same thing will be true for a lot of skills that we want AIsto actually accumulate from all the different workplaces in which they're deployed.
我在别处解释过,为什么我认为真正的持续学习(continual learning,指模型在实际干活的过程中不断把经验写回自己的权重里)是必需的。我不认为,如果 AI 只能靠在一次次会话之间写 markdown 文件来传递经验,它们就能像人一样胜任一整份工作。举个形象的例子:想象学生们就是这样学吹萨克斯的。第一个学生此前从没碰过萨克斯,他走进琴房,试着吹——当然这是他第一次,所以他失败了。他把哪里出了问题写成一堆笔记。门外还等着下一个学生,他进来,把这些笔记全读一遍;他也从没吹过,所以当然也搞砸了,然后继续往这份笔记上加内容。门外排着无穷无尽的学生,一个接一个地给下一个人写笔记。我不认为存在任何一套他们能互相传递的文字,能让后面某个学生第一次上手就把萨克斯吹好。到了某个点,你必须真的把相关经验积累进自己的大脑里。我认为,对于我们希望 AI 从它被部署的各种工作场所里真正积累起来的许多技能,情况也是一样的。
[0:57] Dwarkesh Patel
Okay, so what changes once we have actual continual learning?One, I think that a lot of proposals that have been put forward about regulating AIassume that you train a model and then you deploy it.And therefore, if you run a bunch of checks on the model before it is deployed,we can make sure that it's not going to aid in cyber attacks or do something crazy.I don't think this assumption necessarily makes sense in the future.And this is one of the many reasons I'm actually kind of worried about locking insome kind of safety regulatory regime right now because we don't know what kind of technologywe're going to be dealing with even within a year, let alone within five years or 10 years.What if the model is improving every single day based on the millions of sessions of workit does in that day?
好,那么一旦我们有了真正的持续学习,什么会发生变化?第一,我认为目前提出的很多 AI 监管方案,都假设你先训练一个模型、然后再部署它;因此只要在部署前对模型跑一批检查,我们就能确保它不会协助网络攻击、或者干出什么疯狂的事。我认为这个假设在未来未必成立。这也是我为什么其实挺担心现在就把某种安全监管体制锁死的众多原因之一——我们根本不知道一年之后要面对的是什么样的技术,更别说五年、十年之后。如果模型每天都在根据它当天完成的数百万次工作会话而变得更好,那怎么办?
[1:37] Dwarkesh Patel
If that happens, we could potentially be locking in an archaic and potentially counterproductiveapproach to dealing with the threats from AI.To the extent the government wants some way to do some kind of safety evaluation on model providers,I think it would make more sense to do monthly or quarterly risk inspections rather than tryingto single out some special moment that occurs after training is done, but before deploymentbegins, because that will not be a meaningfully distinct category in the future.Two, how the labs do technical alignment would probably totally need to change.Right now, a lot of research is focused on the question of how we make sure that a frozenset of weights behaves well during deployment.But I'm not aware of much research on the question of how we make it so that even with theconstant weight updates, the AI system never falls prey to jailbreaks or changes into adeceptive or evil persona.And if AIs are consolidating learnings between users as well, how do we prevent users frominjecting backdoors or some kind of malicious inclination into the base model?
如果真是那样,我们可能会把一套陈旧、甚至可能起反作用的应对 AI 威胁的办法锁死。如果政府确实想对模型提供方做某种安全评估,我认为更合理的做法是按月或按季度做风险检查(risk inspection),而不是去挑出「训练结束之后、部署开始之前」的那个特殊时刻——因为在未来,那根本不会是一个有意义的、可以单独切出来的阶段。第二,各家实验室做技术对齐(alignment,让模型的行为符合人类意图)的方式,恐怕得彻底改变。现在大量研究关注的问题是:怎么确保一组被冻结的权重在部署期间行为良好。但我没怎么见到有研究在问:即便权重在不断更新,怎么让 AI 系统始终不被越狱(jailbreak)攻破、不会变成一个欺骗性的或者邪恶的人格?而且,如果 AI 还会把从不同用户身上学到的东西汇总起来,我们要怎么防止用户往基础模型(base model)里注入后门、或者某种恶意倾向?
[2:34] Dwarkesh Patel
In some sense, this is actually kind of what the human alignment problem is, right?Humans improve in a self-directed way.If you have kids, I don't have kids, but I imagine this is what happens.If you have kids, they go out, they learn new things.Sometimes they go crazy.They get one-shotted by crazy ideologies.They take the wrong drug.They become super weird.But you hope that you've given them enough common sense and basic values that they improveas people in a self-directed way without ending up with some super weird beliefs or somemisanthropic ideas.Three, the diversity of AI minds will increase.Right now, there are less than five prominent AI minds, by which I mean the base models, whichare served to millions or hundreds of millions or billions of users at once.And they're all quite similar to each other, by the way, because they've also been all trainedon roughly the same data.But if AIs are learning from experience, and that experience is different between not onlydifferent AI companies, but also between different instances of the same AI model, wecan actually see a lot of diversity come out the other end in this world.And this would be, I think, a net good outcome.I think one of the things to worry about in the future is just having this monolithic
从某种意义上说,这其实就有点像人类的对齐问题,对吧?人类是以自我导向的方式变好的。如果你有孩子——我没有孩子,但我想大概是这样——孩子出门去,学到新东西。有时候他们会走火入魔:被某种疯狂的意识形态一击命中(one-shotted),或者嗑错了药,变得特别古怪。但你希望的是,你已经给了他们足够的常识和基本价值观,让他们能自我导向地成长为更好的人,而不至于最后抱着某些极其古怪的信念或者厌世的想法。第三,AI 心智的多样性会增加。现在,称得上有影响力的 AI 心智不到五个——我指的是那些基础模型,它们同时服务着数百万、数亿甚至数十亿用户。顺便说,它们彼此还相当相似,因为它们大体都是在差不多的数据上训练出来的。但如果 AI 是从经验中学习的,而这份经验不仅在不同 AI 公司之间不同、在同一个 AI 模型的不同实例之间也不同,那我们真有可能在这个世界的另一头看到大量的多样性长出来。我认为这会是一个净收益的结果。我觉得未来值得担心的事情之一,就是最后只剩下一个铁板一块的——
[3:40] Dwarkesh Patel
singleton that's quite boring.A world where we have continual learning would hopefully be more interesting than the modecollapse of different models we see in the world right now.Four, when deployment becomes part of training, the returns to being ahead in the AI race accelerate.Because if you have the best model, and more people are using your AI for more complicatedand useful work.And as a result, they're giving it lots of feedback that can integrate beyond the sessionwindow.Then your model will become even smarter.Five, if the model learns mainly from deployment, then labs will feel a lot of pressure to deploytheir smartest models earlier.Anthropic has reportedly been using Mythos internally since February, but it only shipped the model tothe public in June.In the regime with actual continual learning, this kind of thing would just not be possible.You could not keep a four-month gap between internal and external deployment and still becompetitive because a competitor who ships the worst model on release date will have a smarter modelbased on actual real-world experience.Six, continual learning will create a clear mode for the leading AI labs that they currently lack.Many people have been asking, how will the AI labs actually make money?
——而且相当无聊的独一霸主(singleton)。一个有持续学习的世界,但愿会比我们今天看到的这种「不同模型的模式坍缩(mode collapse,指各家模型的输出风格越来越趋同)」更有意思。第四,当部署本身成为训练的一部分,在 AI 竞赛中领先所带来的回报会加速放大。因为如果你有最好的模型,就会有更多人拿你的 AI 去做更复杂、更有用的工作,结果他们给出的大量反馈能被整合进模型、超越单次会话窗口,那你的模型就会变得更聪明。第五,如果模型主要从部署中学习,那各家实验室会感受到巨大的压力,要更早地把自己最聪明的模型放出去。据报道,Anthropic 从二月起就在内部使用 Mythos,但直到六月才把这个模型交付给公众。在真正的持续学习这个体制下,这种事就不可能了:你没法在内部部署和外部部署之间留出四个月的时间差还保持竞争力——因为一个在发布当天推出的模型比你差的竞争对手,会凭借真实世界里积累的经验,最终拿到一个更聪明的模型。第六,持续学习会给领先的 AI 实验室带来一条它们目前所缺的、清晰的护城河(moat)。很多人一直在问:AI 实验室到底要怎么赚钱?
[4:48] Dwarkesh Patel
I have been asking this.When I had Dario on the podcast, I asked him this question.And he made the analogy of cloud providers.And he made the point, look, the cloud providers are offering many undifferentiated services,but they're earning high profit margins nonetheless.You will have noticed this if you look at Amazon or Google's quarterly earnings.They're doing just fine.But the reason that the cloud margins are so high is that it's really time-consuming andexpensive to switch from one cloud to another.Currently, there's nothing that's stopping me from starting a software repository withcodex and then doing more work on it with cursor and then finishing it up with cloud code.But once we have actual continual learning and the model you're working with is actuallygetting better as it interacts with you from session to session, then there are actuallypretty significant switching costs.If you want to change the AI that you're using, you basically have to fire an employee thathas accumulated months of context on your organization.And you replace them with a very fresh, very unexperienced new intern that you got to retrainfrom scratch.And once you have this kind of lock-in, model providers can demand pretty hefty margins.
我自己也一直在问这个问题。当我请 Dario(Anthropic CEO Dario Amodei)上播客的时候,我问过他。他打了个云服务商的比方,他的意思是:你看,云厂商提供的很多服务是同质化的,但他们照样赚着很高的利润率。你去看亚马逊或谷歌的季度财报就会注意到这一点,他们过得挺好。但云的利润率之所以这么高,是因为从一家云迁到另一家云非常耗时、非常昂贵。而现在,没有任何东西拦着我先用 Codex 起一个软件仓库,然后换 Cursor 接着做,最后再用 Claude Code 收尾。但一旦我们有了真正的持续学习、你手上这个模型确实会随着一次次和你交互而变得更好,那切换成本就相当可观了。如果你想换掉正在用的 AI,你基本上等于要辞掉一个已经积累了好几个月贵公司上下文的员工,再换上一个非常新、非常没经验、你还得从头带起的实习生。而一旦有了这种锁定(lock-in),模型提供方就能要求相当丰厚的利润率了。
[5:50] Dwarkesh Patel
Sorry, I'm a fresh intern.Really, I'm just a fresh intern.Seven.Of course, enterprises will be wise to this kind of dynamic.They will try to avoid this kind of lock-in.But what if the choice is that you either get locked into a model provider or you loseout on this super valuable feature where your model improves for you from session to session?
抱歉,我就是那个新来的实习生。真的,我只是个新来的实习生。第七。当然,企业也会看穿这种动态,他们会设法避免被这样锁定。但如果摆在面前的选择是:要么被某家模型提供方锁定,要么就失去「你的模型会在一次次会话中为你变得更好」这个超级有价值的特性呢?
[6:10] Dwarkesh Patel
If real usage ends up being the main way the models improve, then the AI labs may subsidizeusers and enterprises which allow the model to train on their sessions.This is already happening if you look at the kinds of deals that are offered to new usersof coding products.This is very similar to why Google gives away search.And conversely, the labs may say that any enterprise that refuses to let them train on the sessionscan't have access to the very best models.With both carrots and sticks, the labs can do a lot to get users to allow AIs to learnfrom experience.Now, of course, I'm glossing over the fact that there's a difference between updatingone user's set of weights and pulling all these different weight forks back into themain model.And the latter may be more technically challenging, but in due time, this too will be solved.Eight.AI training already has large economies of scale.You get to amortize all this expensive training across more users.And you see the evidence for this in the fact that the lab revenues are increasing far fasterthan their compute.But continual learning may also lead to economies of scale in inference for end users, namelyfrom batching.You might have seen my episode with Rainer Pope where we discussed this in detail.
如果真实使用最终成为模型变强的主要途径,那 AI 实验室可能会去补贴那些允许模型拿他们的会话做训练的用户和企业。这件事其实已经在发生了——你看看那些面向编程产品新用户的优惠方案就知道。这跟谷歌为什么把搜索白送出去非常相似。反过来,实验室也可能会说:任何拒绝让我们拿它的会话做训练的企业,就别想用上最好的模型。胡萝卜加大棒双管齐下,实验室有很多办法让用户答应让 AI 从经验中学习。当然,我这里略过了一件事:给单个用户更新一份权重,和把这些不同的权重分叉(weight fork)合并回主模型,是两码事。后者在技术上可能更有挑战,但假以时日,这个问题也会被解决。第八。AI 训练本身已经有很大的规模经济了:昂贵的训练成本可以摊薄到更多用户身上。证据就是各家实验室的收入增长远快于它们的算力增长。但持续学习可能还会给终端用户带来推理侧的规模经济,具体来说是来自批处理(batching,把很多请求攒成一批一起算)。你可能看过我和 Reiner Pope 那期节目,我们在里面详细讨论过这个。
[7:13] Dwarkesh Patel
But if per company instructions require full weight updates rather than living in low rankadapters, there's huge advantages from batching.Back of the envelope math suggests that the optimal inference batch sizefor a sparse model like, say, DeepSeq V3 is more than 2,400 concurrent sequences beinggenerated at once.If you don't do this, then you're underutilizing your compute.And if you want to understand why, again, I highly recommend that episode with Raineron inference economics.But anyways, the point here is that a given set of weights is only served efficiently whenthousands of sequences are being decoded against it all at once.A large company with lots of employees and agents who are doing lots of different kindsof things can very efficiently serve their weight fork, whereas an individual user who's onlyrunning a batch size one may suffer more than two orders of magnitude worse efficiency ontheir compute.So the economics of serving personalized weights strongly favor big organizations.Obviously, plenty more will have changed by the time that continual learning actually works.And the most important changes are probably the ones that are hardest to anticipate in advance.But the ones above seem kind of clear even now.
但如果每家公司的专属信息需要走完整的权重更新、而不是待在低秩适配器(low rank adapter,即 LoRA 这类只改一小部分参数的轻量微调)里,那么批处理带来的优势就非常大。粗略估算表明,像 DeepSeek V3 这样的稀疏模型,最优推理批大小(batch size)是同时生成 2,400 条以上的序列。如果你达不到这个量,就是在浪费自己的算力。想搞明白为什么,我还是强烈推荐去看和 Reiner 那期讲推理经济学的节目。总之,这里的要点是:一组给定的权重,只有在成千上万条序列同时对着它解码时,才是被高效利用的。一家有大量员工和 agent、同时在干各种各样事情的大公司,可以非常高效地服务自己那份权重分叉;而一个批大小只有 1 的个人用户,算力效率可能要差上两个数量级以上。所以,服务个性化权重的经济账,强烈偏向大型组织。显然,等到持续学习真正跑通的那一天,改变的事情远不止这些,而且最重要的那些变化,多半恰恰是事先最难预料的。但上面这些,即便在今天看来也已经相当清楚了。
[8:21] Dwarkesh Patel
This was a narration of a blog that I also published on my website.Go check it out at Borkesha.com.Otherwise, I will see you on the next podcast.
以上是我对一篇博客文章的朗读,那篇文章我也发在了自己的网站上,去 dwarkesh.com 看看吧。就说到这儿,我们下一期播客再见。