ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.166 · 全文

Ryan Greenblatt – What happens once AI can automate AI research?

频道: Dwarkesh Podcast
视频: https://www.dwarkesh.com/p/ryan-greenblatt
原文语言: en
统计: 共 158 轮 · Dwarkesh Patel 50 · Ryan Greenblatt 108


[0:00] Dwarkesh Patel

Today I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research,where he focuses on technical AI safety and security work.I want to talk to you about recursive self-improvement.This is the idea that once you build human-level intelligences,they quickly slingshot towards tens of billions of super intelligences,which are each individually more competent than the top human experts across every field.Whether or not this turns out to be the case,I think is actually probably the most important question in the world right now.And historically, I've been quite skeptical that this kind of thing happens,but you seem to think that it might be plausible, and so I wanted to hear the case for it.Yeah, let's talk about this.So first, I think it's worth noting that AIR&D is a type of task at which the AIs are especially good,because both the companies are trying really hard to make their AIs good at AIR&D,and it's the kind of domain, it has a lot of nice properties from the perspective of how AI development works right now.So it's like pretty verifiable.You can do a bunch of stuff iteratively, and he'll climb on various metrics.And then I think once you have AIs, which are roughly matching the top human experts in AIR&D,

今天我要和 Ryan Greenblatt 聊聊。他是 Redwood Research 的首席科学家,主要做技术层面的 AI 安全(safety)与安保(security)工作。我想和你聊的是递归自我改进(recursive self-improvement)。这个想法是说:一旦你造出了人类水平的智能,它们会迅速像弹弓一样弹射到数百亿个超级智能——每一个单独拿出来,都比任何领域里最顶尖的人类专家更强。不管这件事最后是不是真的会发生,我认为这大概是当下世界上最重要的问题。而且从历史上看,我一直对这类事情相当怀疑,但你似乎认为它有可能成立,所以我想听听支持它的理由。

**Ryan:**好,我们就聊这个。首先我觉得值得指出的是,AI 研发(AI R&D)是 AI 特别擅长的一类任务。一方面,各家公司都在拼命让自家的 AI 擅长 AI 研发;另一方面,从当下 AI 开发的运作方式来看,这个领域本身有很多很好的性质。它相当可验证,你可以反复迭代地做很多事情,在各种指标上做爬山式优化。然后我认为,一旦你有了在 AI 研发上大致能匹敌顶尖人类专家的 AI——


[1:02] Ryan Greenblatt

that could sort of kick off a feedback loop where, you know, the AIs are doing AI research,that produces smarter AIs, that feeds back in.And that feedback loop could be strong enough that you end up with a lot of progress in a short period of time.Maybe my sort of median expectation is something like four or five years of AI progress in a single year.And this requires really overcoming a huge amount of diminishing returns in researchand basically doing the equivalent of what progress we would have gotten after a really large compute scale.So this is like a pretty impressive big thing.And it's worth keeping in mind that five years of AI progress, four years of AI progress,even three years of AI progress is really a lot of fucking AI progress, right?

——那就可能启动这么一个反馈回路:AI 去做 AI 研究,研究产出更聪明的 AI,再反馈回去。而这个反馈回路可能强到让你在很短的时间里取得大量进展。我的中位数预期大概是:一年之内取得四到五年的 AI 进展。这需要真正克服研究中巨大的边际收益递减,基本上等于要做出「我们本来得等一次非常大规模的算力扩张之后才能拿到」的那些进展。所以这是一件相当了不起的大事。而且要记住,五年的 AI 进展、四年的 AI 进展,哪怕只是三年的 AI 进展,那也真他妈是一大堆 AI 进展,对吧?


[1:38] Dwarkesh Patel

So, you know, right now it's like three years ago or a little over three years ago,there was GPT-4 that had come out.And right now, of course, we have like, you know, Mythos-5 or whatever,and maybe a somewhat better model than Anthropic has internally.And so that is just a huge amount of progress in a bit over three years.And if we're talking about five years, then maybe we're talking more about like a jump from,you know, GPT-3 to Mythos-5 or whatever.Yeah. Okay. So I think this argument has three different parts,and now I want to evaluate each one of them.First is the argument that AI R&D is very verifiable.Second is the argument that if you automate AI R&D,you could get four or five years of progress in a single year.And third is the argument that what comes out the other end of four or five years of AI progressat the current pace, starting at the current or starting at the starting pointwhenever AI R&D is automated.Yeah.What comes out the other end is an AI where you can drop it on the jobat basically anything you can imagine.You can drop it in Texas politics in the 1940s and it outmaneuvers Lyndon Johnson.You can drop it in, I don't know, TSMC,and it like learns how to do better process engineering at TSMC.

**Ryan:**你看,现在往回数三年、或者说三年多一点,那时 GPT-4 刚出来。而现在我们当然已经有了 Mythos-5 之类的模型,可能 Anthropic 内部还有比这更好一点的东西。所以在三年多一点的时间里,进展量是巨大的。如果我们说的是五年,那可能更接近于从 GPT-3 一跃跳到 Mythos-5 这种量级。

**Dwarkesh:**嗯,好。我觉得这个论证有三个不同的部分,现在我想逐一评估。第一是「AI 研发非常可验证」这个论断。第二是「如果你把 AI 研发自动化,就能在一年内拿到四到五年的进展」。第三是:以当前的速度、从 AI 研发被自动化的那个起点算起,四到五年的 AI 进展之后,从另一头出来的是什么。

**Ryan:**嗯。

**Dwarkesh:**出来的是这样一种 AI:你能把它扔到你能想象到的几乎任何岗位上。你可以把它扔进 1940 年代的得州政坛,它在权谋上能压过 Lyndon Johnson;你可以把它扔进——我不知道——台积电(TSMC),它就能学会在台积电做更好的工艺工程。


[2:49] Dwarkesh Patel

It's certainly a better video editor than I...My video editors are very excellent,but it is just in general better than humans at any given jobthat it finds itself trying to do.So I want to evaluate all of these sub-arguments that lead to basically getting ASI pretty soonafter this benchmark, which you're expecting by 2030 or something, right?

它肯定是个比我……比我的视频剪辑师更好的剪辑师——我的剪辑师们非常优秀——但总的来说,无论它遇到什么工作,它都比人类做得更好。所以我想把这些子论证一个个评估一遍,因为它们合起来基本上是在说:在那个里程碑之后很快就会得到 ASI(超级人工智能)。而那个里程碑,你预期在 2030 年左右,对吧?


[3:10] Ryan Greenblatt

Yeah, I would say that I expect like full automation of AI R&D,perhaps somewhere around like 2031, 2030,and then getting to like the like beats all humans on the job milestone.Maybe I expect median around 2033,but sort of like if I see AI as fully automating R&D,I think I'm expecting that probably within a year.It's just like the way the forecasting works out,there's a...The difference between medians is bigger than the median difference between milestones.Anyway, whatever.By the way, there's this meme on the internetbecause every time I'm trying to ask about people's timelineswhen I'm asking Dario or somebody,I'm always like,okay, how long before getting automated my video editors?

对,我会说我预期 AI 研发的完全自动化大概出现在 2031、2030 年前后;然后到达「在工作上全面击败所有人类」这个里程碑,我的中位数预期大概在 2033 年。不过如果我真的看到 AI 完全自动化了研发,我预期后面这件事大概会在一年之内发生。这只是预测本身算出来的结果:不同里程碑各自中位数之间的差距,比里程碑之间差距的中位数要大。总之,无所谓。

**Dwarkesh:**顺便说一句,网上有个梗,因为每次我问别人的时间表——比如问 Dario 或者别的什么人——我总是问:还有多久我的视频剪辑师会被自动化?


[3:46] Ryan Greenblatt

And there's this meme of like my video editor editing the podcast.Every time I listen to this.The reason I do it is because I think it's easy to get lost inabstractions when you talk about jobs you don't understand welland to very concretely understand what it takes to automate a jobthat I actually understand why it's difficult for LLMs to currently take control over.I do think that the milestone for automating your video editoris earlier than the milestone of being able to automate all human jobs,including like, you know, Texas politics spinning up on the job.So I think, I do think that the video editor automationmaybe occurs more like around full automation of AR&D,but it's very sensitive to how much people are really focusing on understanding video.Yeah.Okay.So let's start with the claim that AR&D is very verifiable.Yeah.So there's a few different parts of this.One of them is that we can train on a bunch of environments,which are like basically directly training the model to do some AR&D taskor some very close by task.So for example, we can have some environment where the model is training some AIon just like eight H100s or whatever, or like some small amount of compute.And that model could be like, you know, the equivalent of like GPT-2 medium or whatever.

**Dwarkesh:**于是就有了那个梗——我的视频剪辑师在剪这期播客,每次我听到都……我之所以老这么问,是因为我觉得当你谈论自己并不真正了解的职业时,很容易迷失在抽象里;而我想非常具体地理解:把一份我确实了解、也确实清楚为什么当前的 LLM 难以接手的工作自动化掉,到底需要什么。

**Ryan:**我确实认为,「把你的视频剪辑师自动化」这个里程碑,会早于「能自动化所有人类工作(包括得州政坛那种边干边上手的活)」这个里程碑。所以我确实认为,视频剪辑的自动化可能更接近于 AI 研发完全自动化的那个时点,但这非常取决于人们究竟在「理解视频」这件事上投入了多少心力。

**Dwarkesh:**嗯,好。那我们先从「AI 研发非常可验证」这个论断开始。

**Ryan:**好。这里面有几个不同的部分。其中一个是:我们可以在一大堆环境上做训练,这些环境基本上就是直接训练模型去做某个 AI 研发任务,或者非常接近的任务。举个例子,我们可以设一个环境,让模型在大概八块 H100、或者说很少量的算力上训练某个 AI,而被训的那个模型可能相当于 GPT-2 medium 之类的东西。


[4:51] Ryan Greenblatt

And then, you know, similar to like nano GPT-2 medium runs or whatever.And in RL, it's like tweaking and iterating on that.And we could do that for a bunch of different tasks.Like we could have it train like image classification models,video generation models, image generation models,all kinds of different sort of ML training tasks.And we could RL it on the task of training increasingly good modelsand also doing things like,oh, here's a particular direction you could pursue for an algorithm.Can you go and implement that?

然后,类似于 nano GPT-2 medium 那种训练跑。而在强化学习(RL)里,就是不断微调、不断迭代。我们可以针对一大堆不同的任务这么干:可以让它去训练图像分类模型、视频生成模型、图像生成模型,各种各样的机器学习训练任务。我们可以在「训练出越来越好的模型」这个任务上对它做 RL,也可以做这样的事:喏,这里有一个你可以尝试的算法方向,你能去把它实现出来吗?


[5:19] Dwarkesh Patel

And so basically there's a whole class of containerizable,verifiable, small scale AR&D tasks that we can aggressively RL the AIs on.And I would say that already companies are presumablydoing some RL on these sorts of tasks.And you could just keep scaling that up,keep making more of these sort of small scale AR&D tasks.And then the AIs could, you know, keep getting better at this.And then implicitly, I'm claiming this will transferto extremely load-bearing aspects of AR&D.But maybe let's stop there for a second and let me get to that part.So let's talk through what this concretely looks like.So you can imagine that we have GPT-7.5.And we say, GPT-7.5, we want to make you so good at AR&Dthat you help us train GPT-9.Okay, so now we want to train GPT-7.5.And we can come up with a bunch of different environments.Like, as you mentioned, we could do,there's already this repo that is the descendantof Andre Karpathy's nano GPT speedrun,where you just try to change everything about the model,from like the optimizer to the hyperparametersto the architecture,to get it to get to a fixed training loss as fast as possible.You could have other kinds of environments,where you could say,hey, GPT-7.5, I want you to train

**Ryan:**所以基本上,存在一整类可容器化、可验证、小规模的 AI 研发任务,我们可以在上面对 AI 做激进的 RL。我会说,各家公司现在多半已经在这类任务上做一些 RL 了。你完全可以继续把它扩大规模,继续造出更多这类小规模的 AI 研发任务,而 AI 就能在这上面持续变强。然后我隐含地在主张:这些能力会迁移到 AI 研发中那些极其吃重(load-bearing)的环节上。不过我们先在这儿停一下,那部分等下再说。

**Dwarkesh:**那我们把这件事具体长什么样过一遍。你可以设想我们有了 GPT-7.5,我们对它说:GPT-7.5,我们想让你在 AI 研发上强到能帮我们训出 GPT-9。好,那现在我们要来训 GPT-7.5。我们可以设计出一大堆不同的环境。比如你刚提到的那种——现在已经有一个仓库,是 Andrej Karpathy 的 nano GPT speedrun 的后续版本——你要做的就是把模型的一切都改一遍,从优化器到超参数再到架构,目标是尽可能快地把训练损失降到某个固定值。你还可以有别的环境,比如你说:嘿,GPT-7.5,我要你训练——


[6:25] Dwarkesh Patel

a really good video game playing model.And I want you to train a model that actually improvesas it plays the same video game again and again.So you learn how to maybe help the modelget better at online learning.Maybe it gets,we don't care how you figure this out.Maybe it's some kind of crazy neural easeor vector memory.Maybe it's some crazy,maybe just like better long context stuff.We don't care.Figure out how to like do online learning research.Obviously, then the fact that GPT-7.5will already have become very good at normal,like it'll be a smart model.And in the same way,the models currently are getting smarter.It'll be better and better at codingin the way that models are currentlygetting better in coding.And you can imagine a hundred other environments like this,which are incentivizing the ability to do AI R&Dby getting GPT-7.5 to like,containerized versions of getting GPT-7.5to develop GPT-2 size models, et cetera, et cetera.And you basically,then you've like,you put GPT-7.5 through a bunch of this kind of training.You build GPT-8.And GPT-8 is now an amazing ML researcher.It has so much intuitionfrom doing all this kind of training.Honestly, a huge intuition pump for meis seeing the progress that AI has made in mathematics,

——一个非常擅长玩电子游戏的模型。而且我要你训练出一个能在反复玩同一个游戏的过程中真的变强的模型。也就是让你想办法帮模型在在线学习(online learning)上变好。也许它靠的是某种疯狂的神经语(neuralese),或者向量记忆;也许是某种疯狂的、就是更好的长上下文机制。我们不在乎,你自己想办法把在线学习这个研究问题解决掉。

显然,到那时 GPT-7.5 本来就已经在常规能力上很强了——它会是个聪明的模型。同样地,就像现在的模型在编程上越来越强一样,它在编程上也会越来越强。

你可以设想另外一百个这样的环境,它们都在激励「做 AI 研发」这个能力:让 GPT-7.5 去做那些容器化版本的任务,比如让 GPT-7.5 去开发 GPT-2 规模的模型,等等等等。然后基本上,你让 GPT-7.5 过了一大轮这种训练,你造出了 GPT-8。GPT-8 现在就是一个惊人的机器学习研究员,它从所有这些训练里积累了大量直觉。

老实说,对我来说一个巨大的直觉助推器,是看到 AI 在数学上取得的进展——


[7:29] Dwarkesh Patel

where I'm just like,if it's a very verifiable domain,AIs can get,even,like mathematics also involves so much,like,I don't really know this object level detailsof the mathematics research,but I'm just like,no, it works.Like,it can just come in like a floodif you can totally put it into a verification loopand it can actually make new breakthroughs.I am curious if ML researchhas a quality of mathematical researchor it seemed like there was a big overhangfrom connecting different disciplines togetheror ideas that were not,no one person would have known enoughabout algebraic geometryand what was the right word?

——我的感觉就是:只要是一个非常可验证的领域,AI 就能……即便像数学这种也包含了那么多……我其实并不真的了解数学研究在对象层面的细节,但我的感觉就是:这条路是走得通的。只要你能把它完全放进一个验证回路里,进展就会像洪水一样涌进来,而且它真的能做出新的突破。

我好奇的是:机器学习研究是不是也有数学研究那样的特质?还是说,那些突破更像是来自把不同学科连接起来所释放的巨大红利——那些没有任何一个人同时懂得足够多的代数几何和……那个词是什么来着?


[8:04] Ryan Greenblatt

Oh man,I really don't know about the math breakthroughs.When no one person would have known enoughabout topologyand algebraic,whatever,blah, blah, blah,in order to make some counterexampleto a big conjecture.My view is that MLis a less deep domain than math.And so there's less of a thingwhere there's like individual expertswith really deep expertisein some area that they combine,but there's definitely going to be some of that.But then I also think that MLhas some attributesthat make it even more favorablethan mathematics in some waysto, you know,AI training.In particular,there's,you can get a better senseof whether you're succeedingand you can see intermediate progress.So in math,it's often the casethat sort of,there's no easy way to seewhether or not you're close to success.Whereas if your goal is to,for example,get to some training loss,you know,2x faster,you can kind of seewhen you're halfway there.And it tends to be the casethat ML innovationsare very additiveor maybe multiplicativedepending on how you think about it.Where basically you can keep stacking innovationsand usually the innovationsjust sort of,just add togetherand don't interfere with each other,though obviously it's going to depend

**Dwarkesh:**唉,我其实真的不了解那些数学突破。就是那种「没有任何一个人同时懂足够多的拓扑学和代数什么什么、blah blah blah,才能构造出某个大猜想的反例」的情况。

**Ryan:**我的看法是,机器学习是一个比数学更浅的领域。所以在机器学习里,「有一批各自在某个方向上有极深造诣的专家,把各自的专长组合起来」这种事更少——虽然肯定也会有一些。但我同时也认为,机器学习有些属性,在某些方面比数学对 AI 训练来说更有利。特别是,你能更清楚地感知自己是不是在往成功走,你能看见中间进展。在数学里,往往没有简单的办法判断你离成功还有多远;而如果你的目标是比如把训练损失降到某个水平、速度快 2 倍,你多少能看出自己什么时候走到了一半。而且机器学习的创新往往是可加的——或者说看你怎么理解,也可以说是可乘的——基本上你可以不断把创新叠上去,而这些创新通常就是简单地相加、互不干扰,当然这显然要看——


[9:07] Dwarkesh Patel

on the details.And so I think thatin a lot of ways,AIR&D will have properties,you know,quite similar to mathwhere basicallyyou can do small,you can like train on chunks of AIR&Dthat are pretty similar in structureto the problem you actually cared aboutin a very verifiable wayand then that will transfer.And then there's an open questionof exactly how well it will transfer.But I think that the transfer currentlyfor math looks pretty good.And my expectation isthat the transfer for AIR&Dwill look pretty good,but not amazing.So one concern I have isI think even in mathematics,as far as I'm aware,we have not seenvery impressive new theory.We've seen a lot of like impressive,verifiable,specific results.For example,find a counterexampleto this conjecture,but we have not seen likecome up with the idea of topologykinds of levels of thingsor come up with thingslike group theory.And it seems like ML of researchhas elements of both of these things,but the less verifiable thingof like come up with new waysof thinking about the problemwould be harder to induce.So it takes, for example,the idea of scaling laws.Obviously,there is someand verification loopssuch thatyou can train GPT-4 betterif you have the idea

Ryan:——具体细节。所以我认为,在很多方面,AI 研发会具有和数学相当类似的性质:你基本上可以在结构上与你真正关心的问题相当接近的 AI 研发小块任务上,以非常可验证的方式做训练,然后这些能力会迁移过去。至于到底能迁移得多好,是个开放问题。但我认为目前数学上的迁移看起来相当好,我的预期是 AI 研发上的迁移也会相当好,但不会好到惊人。

**Dwarkesh:**我的一个顾虑是:据我所知,即便在数学里,我们也还没看到非常令人印象深刻的新理论。我们看到了很多令人印象深刻的、可验证的具体结果——比如「找出这个猜想的反例」——但我们还没看到「提出拓扑学这个概念」那种层级的东西,或者提出群论这类东西。而机器学习研究看起来这两种成分都有,只是那个不太可验证的部分——比如提出思考问题的全新方式——会更难被诱导出来。举个例子,规模化定律(scaling laws)这个想法。显然,是存在某些……以及验证回路的,使得只要你有那个想法——


[10:17] Dwarkesh Patel

of scaling lawsfrom like 2020.But there is a longerand potentially more compute-ladenand like roadto getting,inducing AIsto be like,okay,I got to think carefullyabout how I shouldbe scaling my parametersand data.What are different kindsof investigationsI could runto understand this?

——也就是 2020 年那会儿的规模化定律的想法——你就能把 GPT-4 训得更好。但如果要诱导 AI 自己想到「好,我得仔细思考该怎么在参数和数据之间分配、怎么扩大规模;为了搞懂这件事,我可以跑哪几类不同的探究实验?」,那条路更长,也可能更吃算力。


[10:34] Ryan Greenblatt

Maybe I can likecome up with the visualizationand like a isoflop analysisor something.But that does seem like hard,that does seemlike a longer verification loopthan just,hey,let's get nano GPTlaws to go down.Yeah,let's talk about this.So first of all,I think in the context of math,the thing I would sayis that the AIscan do the equivalentof like baby'sfirst new theoryor whatever,where like,for example,they can just likeprove interestingconjecturesvia like making connectionsand producingnew understandingof like,oh,there's this like thingthat I,this like constructionthat I foundwhich is pretty interestingor like found this likeway of thinking aboutthe problemthat's a bit different.And we do just see that.It's just that the exampleswe see are not likeas impressive as likefounding the fieldof group theory,but like in part,you know,probably founding the fieldof group theoryis like one of the,you know,it's like among the best,biggest mathematicalaccomplishments of all timeand the AIs just aren't,you know,they're not that goodat math yet.And I think thatfrom my perspective,sort of there's a continuumbetween thatand the things we're seeing nowthat the AIsare continuing to march up.Second,I think ML

**Dwarkesh:**也许我能想出某种可视化,做个 isoflop 分析之类的。但那确实显得很难,那确实是一个比「嘿,我们把 nano GPT 的 loss 降下去」长得多的验证回路。

**Ryan:**好,我们聊聊这个。首先,在数学的语境下,我想说的是:AI 已经能做出相当于「婴儿版的第一个新理论」那种东西——比如它们能通过建立联系、产生新的理解来证明一些有意思的猜想:噢,我找到了这么个东西、这么个构造,挺有意思的;或者说我找到了一种稍微不一样的思考这个问题的方式。这些我们确实已经看到了。只不过我们看到的例子,没有「开创群论这个领域」那么震撼——但部分原因是,开创群论大概算是有史以来最伟大的数学成就之一,而 AI 目前在数学上还没那么强。在我看来,从那种成就到我们现在看到的东西之间是一个连续谱,而 AI 正在这个谱上持续向上推进。

第二,我认为机器学习——


[11:37] Ryan Greenblatt

is a very shallow domainrelative to math.So I think in math,there was much moreof a,you find some truedeep abstractionand then like that,like if you reallyunderstand that thing,which is hard to understand,then you get somewhere.Whereas I feel likethe things that arethe equivalent of thatin MLare really likedumb bullshit.Like I'm like scaling laws.Like,come on guys,we can explain scaling lawsreally quickly.And I think the likedeepest and most importantconcepts in math,for example,don't have the propertyof like you can reallyunderstand the underlyingthing and why it mattersin a very short period of time.But I feel likeone effect will bethat we will have gotten ridof all the low-hanging fruitsby 2030.Like I feel like scaling lawswill have been in likewhat math history,Descartes,you know,finding the Cartesian gridand like doing very basicmathematics was.And then eventuallyif we want to keep making progressin the 2030s,it's going to be likedo whatever bullshitis happening at likethe frontiers of mathematicsright now.Yeah,that could be right.My sense is thatjust like some domainsare structurally differentin terms of how they operateand how much they depend onlike sort of deep abstractions

——相对于数学是一个非常浅的领域。在数学里,更多的是那种「你找到某个真正深刻的抽象,而只要你真的理解了那个很难理解的东西,你就能走到某个地方」。而在机器学习里,与之对应的那些东西,我感觉真的就是些蠢兮兮的破玩意儿。比如规模化定律——拜托,各位,我们能很快就把规模化定律讲清楚。而我认为数学里最深刻、最重要的那些概念,并不具备「你能在很短时间里就真正理解底层机制以及它为什么重要」这个性质。

**Dwarkesh:**但我感觉会有这么一个效应:到 2030 年,我们已经把所有低垂的果实都摘光了。我感觉规模化定律在数学史上的位置,相当于笛卡尔发现笛卡尔坐标系、做那些非常基础的数学的阶段。那么到了 2030 年代,如果我们还想继续取得进展,就得去搞那些现在数学最前沿正在发生的破玩意儿了。

**Ryan:**对,这也可能是对的。我的感觉是,不同领域在运作方式上、在多大程度上依赖那种深刻抽象上,本身就是结构性不同的——


[12:37] Ryan Greenblatt

and like physics and mathare much more on the sideof like being very faron the like sort ofvery deep hard to come upwith ideas side.Whereas I think MLand most other domainsare much more amenableto sort of hill climbing.And that's my senseof how this will goin the future.And even in the regimewhere your AIs are like,you know,having to plow,like it's 2030,they need to likea bunch of low-hanging fruitand research has already happened.They need to likemake further progress.I still suspectthat a bunch of the workwill live more on the sideof like buildingincreasingly complicatedinfrastructure,having really good intuitionabout what the experimentsroughly look like.And so I think I'mprobably less sympatheticto like the like thingthat the AIs will lackis like some deep insightand more sympatheticto like they really needa bunch of like tasteabout in the weeds experimentsthat they currently don't haveand need to havea bunch of intuitionfor like what sorts oftraining approachwould workand what wouldn't workin ways that currentresearchers have.And even in caseswhere there has beensome breakthrough in AI,oftentimes in retrospect,it looks like a big bottleneckto making that breakthrough happenwas sort of getting

——而物理和数学更靠近「想法非常深、非常难想出来」的那一端。相比之下,我认为机器学习和大多数其他领域,要更适合爬山式的渐进优化。这就是我对未来走向的感觉。

即便到了那种状态——你的 AI 得埋头苦干,假设那是 2030 年,研究里一大堆低垂果实已经被摘完了,它们还得继续往前推进——我仍然怀疑,很大一部分工作会更偏向于:搭建越来越复杂的基础设施,以及对实验大致该长什么样有非常好的直觉。

所以我想我大概不太认同「AI 会缺的是某种深刻洞见」这种说法,而更认同「它们真正需要的是一大堆关于泥地里那些具体实验的品味(research taste)——这是它们现在还没有的;以及一大堆直觉:什么样的训练方法会奏效、什么不会——而这种直觉现在的研究员是有的」。而且即便是 AI 领域确实出现过突破的那些案例,事后回看,往往做成那个突破的一大瓶颈,其实是把——


[13:40] Ryan Greenblatt

all of the like micro detailsand mungy intuition right.Like an example of thisis when it comesto like training AIswith to be goodat reasoningand chain of thoughtand doing sort ofRL and chain of thought training,it looks like you probablycould have doneRL and chain of thoughton like GPT-3and gotten kind ofinteresting results on mathif you had really scaled it upand done a good job.But at the timethere was low hanging fruitand also doing a good jobwith that trainingis like kind of likein the weedsand then all the technical implementationand scaling it upand getting the hyper parameters right.And so maybe you candemonstrate everythingon like QN1B or whateverand get some sensethat this whole thingis going to work.But people didn't demonstrate itas early as they could havebecause like,you know,of all of these otherlike mungy detailsand intuitionabout exactly how to tunethe parametersand how to set things up.This is my remaining skepticismhonestly about this storyis just,I am,yeah,I'm not sure I understandwhy if research breakthroughsare so amenable to intelligence,why AI progresshas not been historically fasterthan it could have beenand we had to wait for,as you were saying,like by the time

——所有那些微观细节和琐碎脏乱(mungy)的直觉搞对。举个例子,在训练 AI 擅长推理、擅长思维链(chain of thought),也就是做思维链上的强化学习训练这件事上,看起来你当年很可能在 GPT-3 上就能做思维链 RL,而且只要真的把规模扩上去、把活干好,就能在数学上拿到挺有意思的结果。但在当时,一方面还有低垂的果实可摘;另一方面,把那种训练做好这件事本身就很泥泞——所有的技术实现、扩大规模、把超参数调对。所以也许你能在类似 Qwen 1B 之类的小模型上把整套东西演示出来,让人感觉这条路是走得通的。但人们并没有在他们本来能做到的那么早的时候把它演示出来,就是因为所有这些其他的琐碎细节,以及关于到底该怎么调参、怎么把整套东西搭起来的直觉。

**Dwarkesh:**老实说,这就是我对这个故事仍然存疑的地方。我不太确定我理解:如果研究突破真的这么吃智力,为什么 AI 的进展在历史上没有比实际更快?为什么我们非得等到——就像你刚说的——等到——


[14:48] Dwarkesh Patel

RLVR actually worked,even though you could have done itwith like less compute,we had to waitfor oceans of computeand like gigawatts of computeto be availablebefore people are likedoing this trainingon the trajectoryof like this constant,you know,as compute keep increasing,we make more breakthroughs.I don't know,I feel like there werea lot of AI researchersin the year 2022who were trying to crackreasoningand it was just thatthey were like bottleneckedby the ability to writeinfrastructure codeor like what was happening?

——等到 RLVR 真正跑通的时候?即使你本来用更少的算力就能做出来,我们却不得不等到算力多如海洋、等到吉瓦级的算力可用之后,人们才开始做这种训练。那条轨迹就像是恒定的:随着算力不断增加,我们就不断做出突破。我不知道,我感觉 2022 年有一大堆 AI 研究员在试图攻克推理这件事,难道只是因为他们被写基础设施代码的能力卡住了吗?还是当时到底发生了什么?


[15:13] Ryan Greenblatt

It's a complicated mix, right?So I think that they wouldhave gone fasterif they could like,as soon as they thoughtof an experiment,run that experimentwithout bugs,without bugs being very important.And then I thinkanother part of itis that like being ableto run a lot of experimentsat high computelets you paper over waysin which the way you implementedit isn't quite rightor you didn't havethe right hyperparameters.And so I think computeis just like really helpfulfor doing AI researchand you can like,you know,cover over a lot of things.But that doesn't meanthat massive increasesin laborwouldn't also be helpful,especially if that laborcomes with,you know,among the best intuitionsthat people have in the field.I just think that that's,you know,really helpful.I think another partof my perspective here,which is maybe a bit differentfrom where you're coming from,is that I thinkI'm expectingsomewhat more transferthan you seem to be imagining.And I'm imaginingthese AIs are actuallylike pretty good scientistsin generaland are just like,you know,pretty reasonableat all of that stuff.And just sort ofwhen you were to interactwith them,it's not like there's somelike really hyper-specializedsavant type vibe.

这是一个复杂的混合,对吧?我认为,如果他们能做到「一想出某个实验,就立刻把这个实验跑起来,而且没有 bug」——没有 bug 这点非常重要——他们会推进得更快。

另一部分原因是,能在高算力下跑大量实验,可以帮你把「实现方式不太对」或者「超参数没调对」的地方给盖过去。所以我认为,算力对做 AI 研究确实非常有帮助,你可以用它遮掩掉很多东西。但这并不意味着劳动力的大幅增加就不会同样有帮助——尤其是当这批劳动力还自带这个领域里最好的那些直觉的时候。我就是觉得那真的非常有帮助。

我这边视角里还有一部分,可能和你的出发点有点不同:我预期的迁移程度要比你设想的更高。我设想的是,这些 AI 其实是相当不错的通用科学家,在所有这些事情上都相当靠谱。而且当你真的和它们打交道时,感觉不会像面对某种超级专精的学者症候群(savant)式的存在。


[16:09] Ryan Greenblatt

They're actually just likepretty good at all the stuffin AR&Dand then maybe likeextremely goodat some subdomains, right?So they're likeincredibly superhumanat writing kernels,incredibly superhumanat everythingwith very short feedback loops.And then like,you know,pretty good at allthe other stuffand like,you know,just totally ableto match other people.And like,I think we are seeing this now.Like I would say thatwhen I look at AIs right now,I think it's already the casethat they canpretty competently matchlike humanswho are mediocreat ML researchat doing ML research.It's just that being mediocreat ML researchis not that helpful, right?

它们其实在 AI 研发的所有事情上都相当不错,然后在某些子领域上也许极其强,对吧?比如它们在写 kernel 上强到极其超人类水平,在所有反馈回路很短的事情上都极其超人类水平;然后在其他所有事情上也相当不错,完全能和别人打平。而且我觉得我们现在就已经看到这一点了。我会说,现在我看当下的 AI,它们已经能相当胜任地匹敌那些在机器学习研究上表现平庸的人类了。只不过,在机器学习研究上表现平庸并没有那么有用,对吧?


[16:40] Dwarkesh Patel

Like the thingthat you actually wantare people who are goodat ML research.And so my sense isthe AIs are justimproving at all of these things.Their taste is improving.Their intuition is improving.And it's already the casethat their taste and intuitionis not like,it's not like complete garbage.Yeah.So I want to veryconcretely understandwhat it would look likefor five years of AI progressto happen in one year.Yeah.So suppose we were backand when like GPT-3 is developed.And the idea isnot only thatlike basicallywith the level of computethey've hadback in 2022,you could have trainedif we had automated AI R&Dback then,you couldat the end of that yearhave Mythos.That would be the idea, yes.Including with it,like so Mythostook way more computethan they had back then.But like even with the levelof compute they had back then,not only due to allthe breakthroughs,but they also train Mythoswith their level of compute.And what would be requiredis obviouslylike discoveringall the algorithmic progresssince then.Discovering even more actuallybecause you had to make upfor the fact thatlike Mythos uses,I don't know,what was GPT-3 trained on?

**Ryan:**你真正想要的,是在机器学习研究上很强的人。所以我的感觉是,AI 在所有这些方面都在变好:它们的品味在提升,直觉在提升。而且现在它们的品味和直觉已经不是一无是处的垃圾了。

**Dwarkesh:**嗯。那我想非常具体地理解一下:「一年内发生五年的 AI 进展」到底会是什么样子。那就假设我们回到 GPT-3 被开发出来的时候。这个想法不只是说:以他们 2022 年拥有的算力水平,如果我们当时就已经把 AI 研发自动化了,你在那一年结束时就能拿到 Mythos。

**Ryan:**对,那就是这个想法。

**Dwarkesh:**而且还得连带着——Mythos 用掉的算力远远超过他们当年拥有的算力。但即便只有他们当年那个算力水平,不只是要做出后来所有的突破,他们还得用当年那个算力水平把 Mythos 训出来。

**Ryan:**而这需要的,显然是把从那时到现在的所有算法进步都发现出来。实际上还得发现得更多,因为你得补上这么个事实:Mythos 用掉的算力——我不知道,GPT-3 当年是用多少算力训的来着?


[17:36] Ryan Greenblatt

Like 1E23?We can look it up.But is it plausiblyfour ordersof magnitude more compute?Yeah,I think it'ssomewhat less than that.Let's look this up quickly.Um,so GPT-3 training compute,um,is,yeah,it's like 3E23.Um,my sense is that,um,Mythos is probably about,um,a little over three ooms higher.And so the question is,can you overcomethis thousand X compute gapwhile also,you know,um,beating the model?

**Dwarkesh:**大概 1E23?我们可以查一下。但这有没有可能是四个数量级的算力差距?

**Ryan:**嗯,我觉得比那个稍微少一点。我们快速查一下。嗯,GPT-3 的训练算力是……对,大概是 3E23。我的感觉是,Mythos 大概比它高三个数量级(OOM)多一点。所以问题就是:你能不能在跨越这个一千倍算力差距的同时,还把那个模型给打败?


[18:03] Ryan Greenblatt

So here's a concrete claimthat maybe we should,we should talk about.Like,right now,we would be ableto train a modelwith GPT-3 level computethat matches,um,yeah,what exactly do I think?Um,so GPT-3 was,let's say,about,um,yeah,when was it trained?

那这里有一个具体的说法,也许我们该聊一下。就是说,现在的我们,能不能用 GPT-3 级别的算力训出一个模型,它能匹敌……嗯,我到底是怎么想的来着?嗯,GPT-3 大概是……对,它是什么时候训练的来着?


[18:20] Ryan Greenblatt

So it was trained,um,it was trained in,it was released in 2020.So it was trained six years ago.It's worth noting that GPT-3is maybe a little too far awayor too,too far in the past,but let's go with thisfor a second.So GPT-3 was trained,like,about,you know,uh,six and a half,seven years ago.If we were to train a modelwith GPT-3 level computetoday,how can we,how good would that model be?

它是……它是在……它是 2020 年发布的。所以它是六年前训练的。值得说明的是,GPT-3 可能有点太久远、太过去了,但我们先就用它来说。所以 GPT-3 大概是在六年半、七年前训练的。如果我们今天用 GPT-3 级别的算力训一个模型,那个模型会有多好?


[18:41] Ryan Greenblatt

Um,my understanding is based on,like,how algorithmic progress works,we'd be able to train a modelthat's as good as the best modelwe hadperhaps around three years ago.So I think that right nowwe'd be able to train a versionof GPT-3that's probably somewhat betterthan GPT-4is basically what we'd see,um,probably a,yeah,like a moderate amount better,um,than GPT-4.And I think that's about right.I think that roughly lines upwith what,with how algorithmic progresshas worked.Basically the story would end upbeing that to get five yearsof AI progress,you're probably going to needaround,I would say like,maybe eight years of algorithmicprogress,very roughly,um,which is a lot,a lot of algorithmic progress.But it just turns outthat like,most of the AI progressfrom my perspectivehas come from some mixof like algorithms and data.And you can just keep makinglike,um,I think huge improvementson these thingsand training AIswith less compute.So that,I'm glad you brought that upbecause what has happenedsince GPT-3 or even 3.5till now,right?

嗯,按我的理解,从算法进步的运作方式来看,我们能训出一个大致相当于「我们大概三年前拥有的最好模型」的模型。所以我认为,现在我们能训出一个 GPT-3 的版本,它大概会比 GPT-4 好一些——基本上我们会看到的就是这个,大概比 GPT-4 好上中等幅度。我觉得这个判断大致是对的,也大致对得上算法进步一直以来的实际情况。

基本上,结论会变成:要拿到五年的 AI 进展,你大概需要——我会说,非常粗略地讲——大概八年的算法进步,这是很大很大的算法进步量。但事实就是,在我看来,AI 的大部分进展来自算法和数据的某种混合。而你可以在这些东西上持续做出巨大的改进,用更少的算力训出 AI。

**Dwarkesh:**那我很高兴你把这个提出来,因为从 GPT-3、甚至从 3.5 到现在,究竟发生了什么,对吧?


[19:34] Dwarkesh Patel

Like,why is Mytho so good?Um,obviously we've scaledthe compute,we have better algorithms.A huge thing that's happenedis that we have builta deca-billion dollardata industry,which has systematicallycollected and codified,uh,expert human judgmentacross all kindsof different disciplines,codified in the formof RL environments,codified in the formof SFT traces,that these experts builtto help the modelbetter understandhow do you do codingand like,how do you buildcomplex infrastructure projects?

比如说,Mythos 为什么这么好?显然,我们扩大了算力,我们有了更好的算法。但还发生了一件大事:我们建起了一个上百亿美元级的数据产业,它系统性地采集并编码了各行各业的专家人类判断——编码成 RL 环境的形式,编码成 SFT 轨迹(SFT traces)的形式,这些都是那些专家亲手做出来的,目的是让模型更好地理解:编程该怎么做?复杂的基础设施项目该怎么搭?


[20:04] Ryan Greenblatt

How do you do like law?How do you do whatever,whatever?And I,how are the AIsable to replicatethe effectthat currentlyexpert human judgmentseems to be playingin AI progress?Yeah.So my sense is thatscaling up the amountof effort spenton getting expert human datahas not been hugely importantfor AIR&D in general.So in particular,like,you know,over the last few years,we've been scaling upcompute,scaling up peopleworking at AI companiesand scaling upthe amount of effortspent on data labeling.My sense is thatif you like,sort of removethe last like,two doublingsor whateverof data labeling,that would not makea huge differenceor data generationthat would not,I'm sorry,I should say data generationfrom expert humans,that would not makea huge difference.I think a lotof what's been going onis people have beendeveloping better waysto leveragelike humansand AIsto like constructRL environmentsand going somewherefrom that.Wait,how do you explainwhy the AIs have gottenso good at coding?

**Dwarkesh:**法律该怎么做?各种各样的东西该怎么做?那我的问题是:AI 要怎么复现「专家人类判断目前在 AI 进展中似乎正在起的那个作用」?

**Ryan:**嗯。我的感觉是,就 AI 研发整体而言,在「获取专家人类数据」上投入的努力量的扩张,并没有起到特别重要的作用。具体来说,过去这几年我们一直在扩大算力,扩大 AI 公司的员工规模,也在扩大数据标注上投入的努力。我的感觉是,如果你把数据标注——或者说数据生成,抱歉,我应该说是「来自专家人类的数据生成」——的最后两轮翻倍拿掉,那不会造成多大差别。我认为真正发生的很大一部分,是人们在开发更好的方式,去调动人类和 AI 来构建 RL 环境,并在此基础上继续往前走。

**Dwarkesh:**等一下,那你怎么解释 AI 在编程上变得这么强?


[21:02] Ryan Greenblatt

I feel like a big partof that is dataand RL environments,which are likequantifying human experts.But the question is,what is the limiting factoron creating RL environments?My sense of the limiting factoron creating RL environmentswas not so muchlike,like scaling upor like the thingthat drove,the reason whyRL environments todayare much betterthan they werein like,you know,2024is not that muchbecause we have hiredway more human expertsto make RL environmentsand is insteadmuch more becausewe better knowwhat RL,like what RL environmentswe even want to makeand like how we shouldstructure them.And also,we're using huge amountsof AI laborto build RL environments.And I think those effectsare much more importantthan the effect ofhuman laborbuilding the RL environments.I'm not saying

我觉得这里面很大一部分确实是数据和 RL 环境,而这些东西本质上就是在把人类专家量化下来。但问题是:制造 RL 环境的限制因素是什么?我的感觉是,制造 RL 环境的限制因素,与其说是规模上的扩张——也就是说,今天的 RL 环境之所以比 2024 年那会儿好得多,主要原因并不是我们雇了多得多的人类专家来做 RL 环境,而更多是因为我们更清楚自己究竟想做什么样的 RL 环境、该怎么组织它们的结构。同时,我们也在用海量的 AI 劳动力来构建 RL 环境。我认为这些因素的作用,比「人类劳动力构建 RL 环境」的作用要重要得多。我不是说——


[21:49] Ryan Greenblatt

that the human labordoesn't matter.I'm just sayingthere's other,there's other big driversthat are important here.Yeah,I could try to,I could try to argue for this.I mean,one,one thing is just likethe amount ofenvironments people wantor just like,they're very,it's a very large amount.And I think the AIsare actually pretty goodat the taskof making RL environmentsgiven some,some sense of whatthe,the thing should be.There's pre-existing datayou could use.I don't know.A lot of these thingshave good verification loops.If I just look at,for example,this was reportedin Business Insideryesterday thatGoogle is payinglike close to$2 billion for Mechanize.Yeah.Like the,we can just look atmarket ratesor what people thinkreally good human expertsmaking,like human expert datais worth.And it just seems to belike the,the frontier labsseem to thinkit's worth a lot.Yeah.what fraction of,of frontier lab spendingdo you thinkis on datarather than compute?

——人类劳动力不重要,我只是说这里还有别的、别的重要的大驱动因素。

对,我可以试着为这一点辩护一下。首先,光是人们想要的环境数量本身,就非常、非常庞大。而且我认为,只要给 AI 一点关于「这东西大概该是什么样」的感觉,它们其实已经相当擅长「制造 RL 环境」这个任务了。有现成的数据可以用。我不知道,这里面很多事情都有很好的验证回路。

**Dwarkesh:**举个例子,昨天《Business Insider》报道说,Google 为 Mechanize 付了接近 20 亿美元。

**Ryan:**嗯。

**Dwarkesh:**我们完全可以看看市场价,看看人们认为真正优秀的人类专家、人类专家数据值多少钱。看起来前沿实验室是觉得它非常值钱的。那你觉得,前沿实验室的支出里,有多大比例是花在数据而不是算力上?


[22:41] Dwarkesh Patel

Like,what do you thinkis the compute dataspend split?I think it's mostlike overwhelmingly compute,but I also thinkit's because like computeis easier to scale upthan data.But that's really relevantto what's driving progress,right?

就是说,你觉得算力和数据的支出比例是多少?

**Ryan:**我认为绝大部分压倒性地是算力。但我同时也认为,这是因为算力比数据更容易扩大规模。

**Dwarkesh:**但这一点其实非常关系到「究竟是什么在驱动进展」,对吧?


[22:51] Ryan Greenblatt

It's like,suppose,like I agree that,yeah,like my,my sense is that the splitis something likeI would have guessedlike 20 to oneor something,10 to one.I don't know exactly.It depends on the company.I mean,but this is similar to likeoil is 1.5% of GDP,but that mean,but that doesn't meanif you cut oil out,you could like,GDP could continue to run.but it contradictsyour argument, right?

就好比说,假设……我同意,是的,我的感觉是这个比例大概是——我原本会猜大概 20 比 1 之类的,或者 10 比 1。具体多少我不确定,这取决于是哪家公司。不过这就类似于:石油只占 GDP 的 1.5%,但这并不意味着你把石油抽掉之后,GDP 还能照常运转。

**Dwarkesh:**可这跟你的论证是矛盾的,对吧?


[23:08] Dwarkesh Patel

immediately if likeoil went away.Sure,but you,you are just arguingthat because of thehigh market cap,we can learn thatthis is the key driverand I'm sayingthat's not clearly true,right?Sure.Because like you,I think that argumentjust implies,looks,makes it look likecompute is a muchmore important driveror like hiring employeesis a much more important driver.Maybe let's be more concrete.Here's what I,here's what I thinkjust the same waysin my claimis that in GP,if you went back to 2022and you had GP 3.5and you were liketrying to make it betterat codingwithout human experts,I think it would have justbeen very, very difficult.Let me give you an exampleof what I imaginewould be the difficultyfrom going fromGPT-8 to ASI.So one of the thingsyou'd need GPT-8to be good ator like you'd want ASIto be good at is likeI'm going to liketake over a companyand like make itmuch more profitableand like do all kindsof crazy shitto make it work better.I'm going to liketake over a faband like produce more chips.This is like the tier of datathat will,I'm going to likego into Congressand try to convince themto pass some bill,blah, blah, blah.Yeah.This is what I imaginefive more yearsof AI progressat this pace

……如果石油说没就没了,那立刻就出事。

**Ryan:**当然,但你不过是在论证:因为市值这么高,我们就能得出这才是关键驱动力;而我说的是,这一点并不见得成立,对吧?

**Dwarkesh:**当然。

**Ryan:**因为我觉得那个论证反而会推出、会让人觉得算力才是重要得多的驱动力,或者说招聘员工才是重要得多的驱动力。也许我们说得再具体点。我的主张是这样的:如果你回到 2022 年,手上是 GPT-3.5,你想在没有人类专家的情况下把它的编程能力做上去,我认为那会非常非常困难。

**Dwarkesh:**我来举个例子,说说我设想的、从 GPT-8 走到 ASI 的困难在哪。你需要 GPT-8 擅长、或者说你希望 ASI 擅长的事情之一是这样的:我要去接管一家公司,把它的利润做得高得多,为了让它运转得更好,什么疯狂的事都干得出来。我要去接管一座晶圆厂(fab),造出更多芯片。这就是那一档次的数据会……我要走进国会,去说服他们通过某项法案,等等等等。

**Ryan:**嗯。

**Dwarkesh:**这就是我设想的、按现在这个速度再走五年的 AI 进展


[24:09] Ryan Greenblatt

would enable an AIto be able to do.This is the thingI'm really worried about, right?Like the ASIthat can like understandhow to do crazy shitin the world,like what Kissinger can do,can do what likeSteve Jobs can do,et cetera,and also his engineersand stuff.And I'm not surehow you get thatwithout the relevantworld data,which is the equivalentof Mythosbeing really good at codingwhile not havingthe coding environmentsthat have improvedit relative to GPT-3.Yeah.So here are a few points.So first,I bet if you lookat sort of randomly sampledtraining environmentsfor Mythos,they're actually very differentfrom what it looks liketo actually usethe model in practice.My sense is thatthe RL distributionhas like really large deviationsfrom the real worlddata distributionand it's significantlybeing sort of likesmoothed overby a mix of transferand having a small amountof data focusedon the real world.And so my senseis that this will bea similar mechanismas how it worksfor like the,you know,crazy, wildly,like quite superhuman AIyou get as a resultof five yearsof AI progresson top of fullyautomated AR&D.So let's just likego through thisa little bit.So in particular,I think that you couldtrain an AIto be really,

Dwarkesh:……所能让一个 AI 做到的事。这才是我真正担心的,对吧?那种 ASI 能理解怎么在现实世界里干那些疯狂的事,能干基辛格(Kissinger)能干的事,能干史蒂夫·乔布斯(Steve Jobs)能干的事,等等,还有他手下工程师干的那些事。而我不确定,在没有相应的现实世界数据的情况下,你怎么能得到这种东西——这就相当于说,Mythos 在没有那些让它相对 GPT-3 变强的编程环境的情况下,还能把编程做得非常好。

**Ryan:**是的。我这里有几点。第一,我打赌,如果你去随机抽样看看 Mythos 的训练环境,它们其实跟实际使用这个模型时的样子非常不一样。我的感觉是,RL(强化学习)的数据分布和真实世界的数据分布之间存在非常大的偏离,而这一点在很大程度上被两样东西抹平了:一是迁移,二是有少量聚焦真实世界的数据。所以我的感觉是,这跟你在完全自动化的 AI 研发(AI R&D)之上再走五年 AI 进展之后所得到的那种疯狂的、相当程度上超人类水平的 AI,其运作机制是类似的。那我们就把这个过程稍微过一遍。具体来说,我认为你可以训练出一个 AI,让它非常、


[25:16] Ryan Greenblatt

really goodat learning on the flyand doing somethinganalogous toin-context learningbut potentially usingsomewhat different mechanismsin a wide varietyof RL environments.So you buildall these differentRL environmentswhere the AI has tolike adapt on the fly,learn on the fly,figure out what it should do,understand its situation betterand like learn really quicklyfrom feedbackin order to succeedat its objectiveand has thingslike limited resourcesand if it like messes upit can like end upin a much worse position.And then if you trainon a huge numberof these environmentsyou will learnsort of general skillsof like picking upcontext on the flyand we're already seeing this.Like it's already the casethat AIs are nowmuch better at sort ofunderstanding roughlywhat's going onand like picking up contextfrom a, you know,limited amount of informationthey're given access to.And then those AIscould then be puton the job at TSMCand then even thoughTSMC is not like literallyin their data distributiontheir data distributionis really wideand the AIs are extremely goodon their data distributionsuch that it transfersto picking up being goodat, you know,being an engineer at TSMCand learning that on the fly

非常擅长即时学习(learning on the fly),做一些类似于上下文内学习(in-context learning)的事情,但可能用的是有些不同的机制——而且是在种类极其广泛的 RL 环境里训练出来的。所以你搭出一大堆不同的 RL 环境,在里面 AI 必须即时适应、即时学习,弄清楚自己该做什么,更好地理解自己所处的处境,并且要从反馈里飞快地学习,才能完成它的目标;环境里还设置了资源有限之类的条件,而且它一旦搞砸,处境就会变得糟糕得多。然后,如果你在海量这样的环境上训练,AI 就会学到那种通用技能:即时把上下文捡起来。而且我们已经在看到这一点了。现在的 AI 确实已经强得多了——能大致理解正在发生什么,能从给它的有限信息里把上下文捡起来。然后这些 AI 就可以被放到台积电(TSMC)的岗位上去干活。即便台积电并不是字面意义上落在它们的数据分布里,但它们的数据分布非常宽,而且 AI 在自己的数据分布上极其强,强到足以迁移过去,变得擅长当一名台积电的工程师,并且即时把这份工作学会。


[26:18] Ryan Greenblatt

where it looks more likethe way the AI gets goodat being a TSMC engineerisn't that it hasa ton of cash knowledgeon being a good TSMC engineer.It's that it likedoes the equivalentof like some scaled up versionof in-context learningthere.But that'd be the mostprosaic story.Obviously there's likea bunch of different waysthis could go.I think this maybecomes down to thena difference of intuitionabout how far you can get.When I think aboutreally smart people I knowthey're just likenot that effectivein domainsthey don't understand that well.But how longhave they had to learn?

这里更像是这样一幅图景:AI 之所以变得擅长做台积电的工程师,并不是因为它掌握了大量关于「怎么当好台积电工程师」的现成知识,而是因为它在那儿做的是相当于某种放大版的上下文内学习。不过这只是最平淡无奇的一种设想。显然,这件事还可能有一堆别的走法。

**Dwarkesh:**我觉得这最终也许归结为一个直觉上的分歧:你到底能走多远。我想到我认识的那些真正聪明的人,他们在自己不太懂的领域里就是没那么有效。

**Ryan:**可他们有多长时间去学呢?


[26:47] Ryan Greenblatt

No, I agree thatif they had experiencethey would be much better.But that's maybewhat I'm arguing foris that experience of data.Like for exampleif I just get a really smartI don't knowIvy League college gradand I'm likeokay you're now in chargeof negotiatingthe Iran deal.I think they just likewouldn't know what to do.I think if you gotinstead got someonewho is really goodat quickly picking upa bunch of different domainsand you gave themsome time to sort oftrain and talk to peopleand show up their expertiseand do some practicethey would actually dolike a pretty good job.I think most domainsare fundamentally pretty shallowwhere likea very smart generalistwho's good atlike a limited subsetof core skillscan like get goingpretty quickly.And my sense is thatlike that's not truefor literally every domainand my sense is thatthe AIs will developincreasingly good mechanismsfor quickly acquiringunderstanding and expertisein a given domain.So consider for examplehow fast AIs can likeunderstand a new code base.AIs can understanda new code basemuch faster than humans canbut to a degreethat's shallowerthan humans couldcurrently understandbut is getting betterover time, right?

**Dwarkesh:**不,我同意,如果他们有经验,就会强得多。但这也许正是我在论证的东西——那种经验性的数据。比如说,我随便找一个非常聪明的常春藤名校毕业生,然后对他说:好,现在伊朗核协议的谈判归你负责了。我觉得他就是会不知所措。

**Ryan:**我认为,如果你换成找一个非常擅长快速上手各种不同领域的人,再给他一点时间去做些训练、跟人聊聊、把自己的专业度补起来、做点练习,他实际上会干得相当不错。我认为大多数领域从根本上说都相当浅——一个非常聪明的通才,只要在一小部分核心技能上够强,就能相当快地把活干起来。而我的感觉是,这并不是对字面意义上的每一个领域都成立;同时我的感觉是,AI 会发展出越来越好的机制,用来在某个给定领域里快速获取理解与专业能力。比如说,你想想 AI 理解一个新代码库能有多快。AI 理解一个新代码库要比人类快得多,只不过理解的深度还不如人类目前能达到的水平,但这一点在随时间变好,对吧?


[27:48] Ryan Greenblatt

So let me spellthat argument outa bit more.So let's say you takeyou know Fable 5or Mythos 5or whateverand you likewanted to makesome kind of complicatedchange to a reallymassive code basethe model will getsome understandingof the code basevery fastlike in the courseof maybe likeyou knowsignificantly lessthan an hourpotentially muchless than an hourand then its understandingof the code basewill like plateaua little bitwhere it won't getas deep of an understandingas a human would have gottenover a much longer period.So it's sort of likean AI in an hourcan match a humanwith a few weeks maybedepending on the detailsof exactly how complicatedthe code base isbut then it won't matcha human with likeyou knowwho's been workingon that code basefor like two yearsor whateverbut over timethe like amountof understandingAIs can matchhas gone up, right?

那我把这个论证再展开讲一点。假设你拿 Fable 5 或者 Mythos 5 或者别的什么模型,你想对一个非常庞大的代码库做某种复杂的改动,模型会非常快地对这个代码库形成一定的理解——大概在明显不到一小时、甚至可能远不到一小时的时间里;然后它对代码库的理解就会有点儿到达平台期,它不会达到一个人类花长得多的时间才能获得的那种深度。所以这有点像:AI 用一小时能顶上一个人类花几周——具体要看那个代码库到底有多复杂;但它比不上一个已经在那个代码库上干了两年之类的人。不过随着时间推移,AI 所能对得上的理解量级一直在往上走,对吧?


[28:32] Ryan Greenblatt

So if we look at like3.7 Sonnetor 3.5 Sonnetmaybe it could only matchthe equivalent ofunderstanding a code basefor like a dayor somethingbut nowyou knowAIs are much betterat like sort ofbuilding contextabout a taskand so you can be likeMythosI want you to reallyunderstand this code baseand then you knowthen implement this featureand it will likespawn a bajillion sub-agentsthose sub-agentswill pour overa bunch of thingsit will likedeliver a bunchof context backit will then likeinvestigate a few thingsand it's not likeamazing at doing thisbut it's likeit can happenlike really fastand it can workpretty welland it's not very hardfor me to imaginehow you could train AIsto be increasingly goodat this task, right?

所以如果我们看 3.7 Sonnet 或者 3.5 Sonnet,也许它当时只能对得上「理解一个代码库一天」的水平;但现在,AI 在为一项任务构建上下文这件事上强得多了。于是你可以对 Mythos 说:我要你真正吃透这个代码库,然后去实现这个功能。它就会派生出无数个子智能体(sub-agents),这些子智能体会把一堆东西翻个遍,把一大堆上下文送回来,它再去查证几件事。它做这件事并不算惊艳,但速度真的很快,而且效果相当好;而且对我来说,要想象怎么把 AI 训练得在这项任务上越来越强,一点都不难,对吧?


[29:07] Ryan Greenblatt

The task of likeimplement some verycomplicated featurein some reasonable wayin a very big code baseis extremely verifiableand that can likebe a thing the AIsimprove onand similarly likethere's a broader scaleof like quicklyunderstanding contextand being able to likehave a bunch of different AIslearn in paralleland then merging that togetherI think there seems to bea crux herewhich I thinkis just an empirical questionwe'll see onwhich ishow good is a transferbetween gettingreally really goodatunderstanding the situationgetting up to speedmaking progressover long periodsin verifiable domainswhich the AIsare obviously gettingway way better atreally fasttookaygo talk to the presidentand like convince himto do X thingor goyou're now in chargeof Googleand now youyou must make Googlea much more profitablecompany this quarterlet me try to justspell out a few morearguments that aremaybe relevantso one thing is thatI do think that likewhen looking at likehow the AIs haveimproved essay writinglet's talk about thata little bitso I think there'sone thing which isthat you can getsome dataeven on these domainsand AIs will be ableto get some dataeven on these domainswhen on a very fastprogress trajectory

「在一个非常大的代码库里、用某种合理的方式实现某个非常复杂的功能」这项任务是极其可验证的,所以这可以成为 AI 不断改进的一个方向。类似地,还有一个更大尺度的版本:快速理解上下文,并且能让一大堆不同的 AI 并行地学,再把学到的东西合并到一起。

**Dwarkesh:**我觉得这里似乎有一个关键分歧点(crux),我认为它就是个经验问题,将来会见分晓,那就是:从「变得非常非常擅长理解处境、快速进入状态、在可验证的领域里于长时段上持续推进」——AI 显然正在这方面飞快地变强——迁移到「好,去跟总统谈谈,说服他做某件事」,或者「现在谷歌归你管了,你必须在这个季度让谷歌的盈利能力大幅提升」,这中间的迁移效果有多好?

**Ryan:**让我再把几个可能相关的论点摆一摆。第一点是,我确实认为,看看 AI 在写文章这件事上是怎么进步的——我们就聊聊这个——我想有一点是:即便在这些领域里,你也能拿到一些数据;而在一条非常快的进展轨迹上,AI 也将能在这些领域里拿到一些数据。


[30:18] Ryan Greenblatt

so like maybe it's hardto build like averifiable environmentfor like was your essayreally goodaccording to humansbut you can doa bit of thatyou know you can dosome trainingyou can do someonline trainingand the AIs will beable to do somelike you knowonline trainingbased on real worldstuffthey'll be ableto like have evalsthey'll be ableto like sample thatand you can you knowscale up the cadenceat which you do thisand then the secondthing is that in practicewhen I just lookat the transferit seems okaylike I think that in factthe AIs have improveda bunch at non-verifiabledomainsand it is in factthe case that it'shard to point to likedomains that arereally hard to verifyon which the amountof improvementbetween you knowGPT-4 and Mythospractice and nowthat doesn't meanthat Mythos is likebetter than the besthumans or somethingright it can stillbe like significantlyworse than typicalhuman professionalsat some aspectof their jobwhile still beinglike way betterthan GPT-4which was likenot even closeyeah so we'retalking about howhow important dataversus algorithmicprogress has beenfor explaining theprogress of the lastfew yearsthat reminds meI'm actually runningan experiment withJerry Han who's

所以,也许很难为「按人类的标准,你这篇文章写得到底好不好」搭出一个可验证的环境,但你能做到一部分。你可以做一些训练,可以做一些在线训练;AI 也将能基于真实世界的东西做一些在线训练,它们能有评测(eval),能去做采样,而且你可以把做这件事的节奏规模化地提上去。第二点是,实际去看迁移效果,我觉得还行——我认为事实上 AI 在非可验证的领域里已经进步了一大截;而且事实上,你很难指出哪些领域是真正难以验证、以至于从 GPT-4 到 Mythos 之间在其上几乎没有进步的。这并不意味着 Mythos 就比最顶尖的人类更强之类的;它在某项工作的某些方面,仍然可能明显不如典型的人类专业人士,但同时又比 GPT-4 强得多——GPT-4 在那些方面压根就不沾边。

**Dwarkesh:**是的,我们刚才在聊的是:在解释过去这几年的进步时,数据和算法进步各自有多重要。这让我想起来,我其实正在和 Jerry Han 一起做一个实验,他


[31:18] Ryan Greenblatt

actually still acollege studentwhat we're basicallydoing to evaluatehow muchhow much progressis coming fromdata versusalgorithmsis training thebest algorithmicrecipe from 2019till nowwith the best datafrom like the2026 data fileand then alsotraining thedifferent data filesgoing back to2019 to 2026with the currentbest training recipelike the algorithmicrecipeyeahand I think thatwill be aninterestingI'm curious ifyou want topre-register likewhat amountof multipliersare coming fromone versus theotherso we need tobe pretty carefulwith what wemean when wesay the worddataso I was tryingto be prettycareful todistinguish betweenscaling upspending ongetting humanexperts tolabel dataor like scalingup the amountof humanexpert labeldatapre-trainingdata doesnot comethe reason whywe have abetter pre-trainingdata set nowversus in 2019is not becausepeople are spendingway more moneygetting humanexperts to liketype up datathat the AIsare then trainedonI think it'spartiallyI think it'snot much of itI think it'svery little ofthe pre-trainingdata improvementsI think thevast majorityof the pre-trainingdata improvementswhich I do meanpre-trainingwe should talkmaybe separatelyabout mid-trainingpost-training

Dwarkesh:……其实还在念大学。我们做这个实验,基本上是为了评估进步里有多少来自数据、多少来自算法:做法是把从 2019 年到现在最好的算法配方,配上大概 2026 年那份数据文件里最好的数据来训练;同时反过来,把 2019 到 2026 各年份的不同数据文件,配上当前最好的训练配方、也就是算法配方来训练。是的,我觉得这会挺有意思的。我想问你要不要先做个预注册(pre-register):这两边各自贡献的倍数是多少?

**Ryan:**那我们得非常小心地界定,说「数据」这个词时到底指什么。我刚才一直很小心地做这个区分:一边是加大投入、花钱请人类专家去标注数据,或者说扩大人类专家标注数据的量;预训练数据不是这么来的。我们今天之所以有比 2019 年更好的预训练数据集,原因并不是人们砸了多得多的钱、请人类专家把数据敲出来给 AI 训练。我认为这里面有一部分是……我认为这部分占比不大,我认为在预训练数据的改进里,这一项占的比例非常小。我认为预训练数据改进里的绝大部分——我这里说的确实是预训练,中训练(mid-training)和后训练我们也许该另外聊——


[32:20] Ryan Greenblatt

but I think thevast majorityof pre-trainingdata improvementsare fromscience onbetter understandingwhat data setsare goodand schleppylabor onfiguring outhow to filterdownand so myview is thatimprovements ofthe form oflike you knowopen web textto fine webor whateverlike thatimprovement isbetter describedas aalgorithmicimprovementof the sortthat you canyou knowstudy withsome GPUsand then doand you don'tneed humansto like youdon't needhuman expertdata to dothatnow there's adifferent effectwhich we couldtalk aboutwhich is thatmaybe theinternet in2026it has muchmoreis more of afertile groundfor trainingdata thanlike theinternet in2018like it's likethere's alsobeen an effectwhere likethere's justmore humansposting on theinternetso there'smore dataharvestmy senseis thatthat effectis goingto bequite a bitsmallerthan theeffect ofjust likehumanslike knowingbetter howto curatethe datahaving betterscrapesknowing howto processthose scrapesbetterthis sortof thingthis is morelike automatedengineeringand automatedR&Dthat's rightthat makessenseso likeI thinkthat insome sensethe thingyou wouldwant tolook atis be likewe're goingto dotwo posttrainingpipelinesone posttrainingpipelinewhere weonly havelike a

但我认为,预训练数据改进的绝大部分来自两样东西:一是科学,也就是更好地搞清楚什么样的数据集才是好数据集;二是苦力活,也就是琢磨怎么把数据过滤下来。所以我的看法是,像从 OpenWebText 到 FineWeb 这一类的改进,更恰当的描述是一种算法改进——那种你用几张 GPU 就能研究、就能做出来的改进,你并不需要人、并不需要人类专家数据就能做到。

现在还有另一个效应,我们也可以聊:也许 2026 年的互联网,比起 2018 年的互联网,是一片更肥沃的训练数据土壤;也就是说,还存在这样一个效应——在互联网上发帖的人变多了,所以能收割到的数据也更多。我的感觉是,这个效应会比另一个效应小相当多:那就是人类更懂得该怎么筛选整理数据、有了更好的爬取、更懂得怎么把爬来的数据处理得更好,诸如此类。

**Dwarkesh:**这更像是自动化工程和自动化研发(R&D)。

**Ryan:**没错。

**Dwarkesh:**这说得通。

**Ryan:**所以我觉得,在某种意义上,你会想看的东西是这样的:我们来跑两条后训练流水线。一条流水线里,我们只有极少量的人类专家来做


[33:22] Ryan Greenblatt

tiny numberof humanexpertsto dothelabelingbut wecan havelike youknowsmartAIsand thenanotherlike youknowyou'relikewe'regoing tobuildMythos 5is goingto builda posttrainingpipelinebut itonly hasaccessto likeinternetdataplus likea tinyamountof humanexpertsmethods

标注,但我们可以有很聪明的 AI;另一条则是——比方说,我们让 Mythos 5 去搭一条后训练流水线,但它只能拿到互联网数据,外加极少量的人类专家。所用的方法


[33:43] Ryan Greenblatt

we hadin 2024but withlikea shit tonof humanexpertsand againboth havethe internetdatamy senseis thatthecurrentmethodsbut withoutmany humanexpertsactually willdo quitewellthough it'sa bitmessybecauselikeMythoslikeit'slikecanMythosgetsomethingthat'smorecapablethanMythoslikeyoumightneedtobea bitthoughtfulonlikewhatmodelisitthatyou'reposttrainingwhatisyourviewonwhatistheleastverifiabletheleastverifiableprobablymakingcallsonlargeexperimentslikethethingthatIthinkismostlikelytobesortofthebottleneckintermsoftheAIsarereallygoodatverifiabledomainsbutnotdoingtheactualthingisjustlikebigexperimentsyouonlygeta fewtrieswella fewismaybea bitunderstatedbutbasicallyhistoricallyR&Dhas beendrivenby doingnear frontierscaleexperimentsandthathasbeenprettyimportantandactuallydoingtheonebigtrainingrunwhereyoudecideexactlywhattoincludeinthatandthere'sa bunchofwaysthattheAIscanmakethatmoreverifiablesotheyhave

是我们 2024 年就有的那一套,但配上他妈的一大堆人类专家;而且两边都有互联网数据。我的感觉是,用当前的方法但没有多少人类专家,实际上会干得相当不错。不过这事儿有点乱,因为 Mythos……问题在于,Mythos 能不能搞出比 Mythos 更强的东西?所以你可能得稍微仔细想想:你要做后训练的到底是哪个模型。

**Dwarkesh:**你怎么看,最不可验证的是什么?

**Ryan:**最不可验证的,大概是对大型实验做判断拍板。我认为最有可能成为瓶颈的地方就在这儿:AI 在可验证的领域里非常强,但真正要干的那件事——也就是跑大实验——你只有很少的几次尝试机会。嗯,「几次」这个说法也许略微低估了;但基本上,从历史上看,研发一直是靠做接近前沿规模的实验驱动的,这一点相当重要;还有真正去跑那一次大的训练——你要精确地决定这次训练里放进什么。而 AI 有一堆办法能把这件事变得更可验证,所以它们


[34:57] Dwarkesh Patel

beenscaleduplessthanyouwouldhaveotherwiseexpectedandlikeforexamplecostofpertokenhasn'tincreasedas muchas youmighthavethoughtisbecausethereisabenefittodoingmoreofyourworkatsmallscalewhereyoucanrunmoretrainingrunsandgetmorecyclesinandsoyou'renotaslikeyouknowleaningashardononebigreallyimportanttrainingrunIIjustwanttounpackacoupleofthingsthatwerefortheaudiencethethingyou'repointingoutisIthinkthepricepertokenhasnotincreasedthatmuchsince20242023yeahsoGPT-4waslikeIdon'tknowlikewaslike$30peroutputtokenandlikeMythosis$50pertokenrightandsothethingyou'retryingtoexplainishowcanitbethatyou're

Ryan:……被放大的程度比你原本预期的要小。举个例子,每 token 的成本没有像你可能以为的那样涨那么多,原因就在于:把更多工作放在小规模上做是有好处的——在小规模上你能跑更多次训练、多转几个循环;这样你就不必那么严重地押注在某一次特别重要的大训练上。

**Dwarkesh:**我想给听众拆解一下刚才那几件事。你指出的是——我理解——每 token 的价格从 2024 年、2023 年以来并没有涨多少。

**Ryan:**是的。

**Dwarkesh:**GPT-4 当时大概是——我不太确定——每输出 token 30 美元,而 Mythos 是每 token 50 美元,对吧。所以你要解释的是:怎么会出现这种情况——你


[36:18] Ryan Greenblatt

ofabustandpartofitisthatIthinkthere'sjustabunchofdetailsandactuallygettingthatrightandsoitmakessensetojustdomoreoftheworkatsmallerscaleandjusteatthefactthatyou'retakingahitonfinalperformanceinordertobeabletoquicklyiterateandtrainmoremodelsfasterandbetterlearnandbetterbeabletohaveasmarterultimateproductionmodelthisisnottheonlyeffectthere'salsothefactthatRLbenefitsmorefromsmallmodelsthere'sa bunchofthingsgoingonbutIdothinkthatin factpeoplearemakingtrade-offstowardsthesideoffasteriterationtimesbecauseofalgorithmicprogressbeingsofastitthese

……成了一场泡汤。其中一部分原因是,我认为这里面就是有一大堆细节,而真正要把这些细节做对(并不容易)。所以合理的做法就是把更多工作放在更小的规模上做,并且认了——你在最终性能上要吃点亏,换来的是能快速迭代、更快地训练更多模型、学得更好,最终能拿出一个更聪明的生产模型。这不是唯一的效应,还有一个事实是 RL 从小模型上获益更多;有一堆因素在同时起作用。但我确实认为,事实上人们正在把权衡往「更快的迭代速度」这一侧倾斜,因为算法进步实在太快了。这些


[37:09] Dwarkesh Patel

kindsofmistakeswheretheymightgetreallygoodatengineeringandbeingtrainedtoavoidbugsbasicallytheoppositeoftheslopworldweliveinnoworarelivinginlessandlessovertimebutthenthere'salsothequestionofcantheydotheanalysistofindtherightexperimenttoidentifywhatisgoingwrongwiththetrainingrunrightnowwhichseemstobeverybottleneckedby thetasteofextremelyfewhumanswhoarelikerightnowmyassumptionisGDMis goingthroughthisrightnowwherehumansaretryingtofigureoutwhatiswrongwiththetrainingpipelineandyeahthere'ssomerumorthatrightafterNoam ShazirjoinedbackorjoinedGDMwhichhe'snowlefttheyhadanewreallygoodtraining!

……这类错误——它们可能会变得非常擅长工程,并被训练成能避免 bug,基本上就是我们如今身处的那个「垃圾产出(slop)世界」的反面,而且我们身处其中的程度随时间在越来越低。但接下来还有一个问题:它们能不能做那种分析,找出正确的实验,去定位当下这次训练到底哪里出了问题?这件事眼下似乎极度受制于极少数几个人的品味。我现在的猜测是,GDM 此刻正在经历这个过程:一群人正在试图搞清楚训练流水线出了什么毛病。还有,有传闻说,就在 Noam Shazeer 回归、或者说加入 GDM 之后不久——他现在已经离开了——他们就有了一次非常成功的训练


[37:50] Ryan Greenblatt

runthathappenedandthereasonwhyisNoamShazirjustlookedattheircodebaseandfoundabunchofbugsbecauseheknewwheretolookmysenseisthattrainingAIstofindbugsisgoingtobeoneoftheeasiertaskstotrainAIsonbecausemostofthesebugswe'retalkingaboutcanprobablybedemonstratedwithoutthatmuchcomputeandprobablyyou getprettygoodtransferfrompointingoutothertypesofbugsatsmallera

Dwarkesh:……跑了出来;原因是 Noam Shazeer 看了一眼他们的代码库,就找出了一堆 bug,因为他知道该往哪儿看。

**Ryan:**我的感觉是,训练 AI 去找 bug,会是训练 AI 做的比较容易的任务之一,因为我们说的这些 bug 大多数很可能不需要太多算力就能复现出来;而且从「在更小规模上指出其他类型的 bug」很可能能获得相当好的迁移。这是一项相当


[38:22] Ryan Greenblatt

prettyverifiabletaskit'snotarbitrarilyverifiablebecausemaybeoftentodemonstratethebugyoumightneedtodoamoderatescalecomputeexperimentwhereyouspinupthewholedistributedinfrastructureandrunitbutoftentimesyou'llbeabletodemonstrateitprettyconvincinglyatsmallerscalein awaywhichyoucouldtrainonandsomysenseisitwouldn'tbesurprisingifrightnowpeoplehaveRLenvironmentswheretheyintroduceasubtlebugintosometrainingrecipetraintheAItopointoutthesubtlebugandthenhavearubricwherethey'relikediditactuallyfindtherightbugandthatseemsverydoableandyoucoulddoabunchofstuffthere'sabunchofthingsyoucoulddoalongtheselinesthatIthings

……可验证的任务。它并不是任意可验证的,因为也许很多时候,为了复现那个 bug,你得做一个中等规模的算力实验,把整套分布式基础设施拉起来跑一遍。但很多时候,你能在更小的规模上、用一种可以拿来训练的方式,相当有说服力地把它复现出来。所以我的感觉是,如果现在就有人在搞这样的 RL 环境,我一点也不会惊讶:他们往某个训练配方里塞进一个隐蔽的 bug,训练 AI 去把这个隐蔽的 bug 指出来,然后设一套评分标准(rubric),看它到底有没有找到那个正确的 bug。这看起来非常可行,而且你还能做一大堆别的事;沿着这条思路能做的东西有一大堆。我觉得有些东西


[39:15] Ryan Greenblatt

thatareanalogoustohyperparametersandthat'sthethingthattheAIsmightmoststrugglewithbutIcurrentlyexpectthere'llbeenoughtransferifyoutrainonallthesedifferentenvironmentsthattheAIswillbegoodatthatdomainandIshouldbeclearIalsothinkthattheAIswilltransfertootherdomainsIthinkthere'sgoing tobethedomainstheAIsareby farthebestatthenthere'sdomainswherethey'resomewhatlessgoodatandIthinkwestillseetransfertoeverythingandit'sreallyhardfor metothinkofexamplesofcognitivetaskshumansdowherewe'renotseeingsometransferfromAIimprovingsolet'sstepbackandpackagethiswholestorysoIthinkpeoplemayprobablyfollowalongwiththestoryofwehaveGPT7.5trainonthat

……类似于超参数(hyperparameters),那可能是 AI 最吃力的地方。但我目前预期,如果你在所有这些不同的环境上训练,迁移会足够多,AI 在那个领域也会变得很强。我还得说清楚,我也认为 AI 会迁移到其他领域。我认为会有这么一个格局:有一批领域是 AI 强得最离谱的;然后有一批领域它们没那么强。但我认为我们仍然会看到向一切领域的迁移;而且我实在很难想出这样的例子——人类会做的某项认知任务,在 AI 变强的过程中我们完全看不到任何迁移。

**Dwarkesh:**那我们退一步,把这整个故事打个包。我想大家大概能跟上这样一个故事:我们有了 GPT-7.5,在那上面训练


[40:04] Ryan Greenblatt

arebetteratplayingvideogamesthatrequiresampleefficiencyoronlinelearningorwhateverothercapabilityanotherthingthat'sreallyimportantisyoudon'tjustdoGPT 2sizedrunsyoualsodosmallfinetuningrunsonGPT 6youhaveGPT 2and youcandofullpre-trainsonGPT 2andthenyoucandosmallposttrainingormidtrainingrunsonGPT 6andthenyoucandoasmallnumberofexperimentsthatareactuallyatfrontierscalebutyoudoabitofonlinetrainingorsomethingwhatdomeanbydoonlinetrainingonthatyeahsoanotherthingthatwecandoiswecantakeGPT 7.5andpresumablyin thecourseofGPT 7.5workit'srunningabunchofexperimentsatvaryingscalethatareonthecriticalpathforAIformanyofthosethingsyou'llbeabletogetasenseafterthefactforwhetherornotitdidagoodjobrightsolikeitdidsomeyouknowposttrainingexperimentwhereitwastryingtofigureoutwhethersomemethodactuallyworksandinsomecasesyou'llbelikewhoaitfoundthislikekickassmethodittotallyworkedandthenyoucanreinforcethatbyjustlikeImeanonethingyoucoulddowouldbeliketakethatbehaviorconverttheexperimentyoujustranintoanRLenvironmentbasedonproductiondataandthentrainonthatoryoucouldpotentiallyjustliterallytaketherolloutsthatfoundthatanddosomesortofoffpolicyRLoryoucoulddosomeonpolicyRL!

Ryan:……在那些需要样本效率、在线学习或者别的什么能力的电子游戏上玩得更好。另一件非常重要的事情是:你不会只跑 GPT-2 那种规模的训练,你还会在 GPT-6 上跑小规模的微调。你有 GPT-2,你可以在 GPT-2 上做完整的预训练;然后你可以在 GPT-6 上做小规模的后训练或者中训练;然后你还可以跑少量真正处在前沿规模上的实验,但只做一点点在线训练之类的东西。

**Dwarkesh:**你说的「在那上面做在线训练」是什么意思?

**Ryan:**是这样,我们还能做的另一件事是:我们可以拿 GPT-7.5——想必在 GPT-7.5 干活的过程中,它会跑一大堆不同规模的实验,而这些实验都处在 AI 的关键路径上。对其中很多事情,你事后是能判断出它干得好不好的,对吧。比如说,它做了某个后训练实验,想搞清楚某种方法到底管不管用;在某些情况下你会说:哇,它找到了一个牛逼的方法,完全奏效。然后你就可以去强化这个行为——我是说,你能做的一件事是:把这个行为、把它刚跑完的那个实验,基于生产数据转化成一个 RL 环境,然后在上面训练;或者你也可以干脆直接拿到发现该方法的那些 rollout,做某种离策略(off-policy)RL;或者你也可以做某种在策略(on-policy)RL。


[41:33] Dwarkesh Patel

butthenitactuallydoesrealR&DinthepracticeoftryingtobecomebetteratAIR&Dandyouarelikethisisacoolthingthatyoudiscoveredlet'sactuallyalsousethisinproductioninthefutureandteachyouhowtouseitinproductionthat'srightsteppingbacksoGPT-8becomesGPT-8as a resultof allthisAIR&DtrainingandjustgenerallybecomingsmarterthenithelpsyoubuildGPT-9andanotherveryimportantthingthatwouldhavehadtohappenwhichismaybethethingI'mmostskepticalofisGPT-8hasfiguredouthowtomakeitsothatwhateverit'sdoingtomakeGPT-9asintelligentasit isitstillneedsthehumanscurrentlylikeAIresearcherstrytheirshitandthey'relikeokaybutwetrainedGPT-4.5or

Ryan:……但接下来,它是在「努力让自己更擅长 AI 研发(AI R&D)」的实践中真正做研发。而你会说:这是你发现的一个很酷的东西,我们以后在生产里也用起来吧,并且教你怎么在生产里用它。

**Dwarkesh:**没错。退一步说:GPT-8 之所以成为 GPT-8,是这一整套 AI 研发训练的结果,加上它整体变得更聪明;然后它帮你造出 GPT-9。但还有一件非常重要、必须发生的事——这可能是我最怀疑的一点——就是 GPT-8 得搞清楚怎么做到这件事:无论它用什么办法把 GPT-9 造得跟它自己一样聪明,这套办法目前还是离不开人类的。就像现在的 AI 研究员们试各种招数,然后说:好吧,可我们训练出的 GPT-4.5 在


[42:22] Dwarkesh Patel

someevaluationoftryingtousethemodelinproductionanditwasn'tthatgoodandwe'renotgoingtoshipitandsoGPT-8needsthisabilitytoseehowgoodthetransferistoalltheseotherthingsyou'retalkingaboutbeingreallygoodatTexaspoliticswhichisnotaproductionenvironmentandinfactcannotbeacontainerizedenvironmentgiventhenatureofthetaskintheagentsgetlongerandlongerhorizontheshorthorizonthingsyou cancontainerizeislikeokaycodethisuporwhateverextremelylonghorizonthingslikegorunasuccessfulbusinessgohaveaprofitabledayinthemarketsgonegotiateatradedealorwhateverthesethingsareactuallyveryhardtocontainerizeandsoIthinkit'sveryplausibleto methatit'sveryhardforGPT-8tofigureouttraining

……某项「试着把模型用在生产里」的评估中表现得没那么好,所以我们不打算发布它。因此 GPT-8 需要具备这样一种能力:看清它到底能多好地迁移到你说的所有那些别的事情上——比如非常擅长得克萨斯州的政治,而那并不是一个生产环境;事实上,考虑到这类任务的性质,它也没法做成一个容器化的环境。随着智能体的任务时长(time horizon)越来越长——短时长的事情你可以容器化,比如「好,把这段代码写出来」之类;但极长时长的事情,比如「去把一门生意做成」「去在市场上做出盈利的一天」「去谈成一笔贸易协议」之类,这些东西实际上非常难容器化。所以在我看来,很有可能出现这种情况:GPT-8 很难搞清楚,训练


[43:15] Ryan Greenblatt

doesn'tgeneralizeinthatwayyeahsoaconcernyoumighthaveislikewetrainGPT-8andGPT-8justlikeisagainbetteratalltheR&DtasksthatwecanmeasurebutisnotgoodattheyouknowsomedownstreamtaskswecareaboutsoIthinkIhaveafewpointssofirstIthinkit'slikeIkindofIexpectthatifyoudotheobviousthingyoudogetprettygoodtransferandyou'llbeabletoholdoutsomeoftheobviousstuffyou'redoingandwhenIsaydotheobviousthingIjustmeantrainonawidevarietyofenvironmentswheretheAIhastoaccomplishweirdobjectivesandlearnaboutwhat'sgoingonandIthinkyou'llbeabletogetsomefeedbackwithsomeenvironments!

Dwarkesh:……并不会以那种方式泛化。

**Ryan:**是的。所以你可能会有这样一个担忧:我们训练出 GPT-8,而 GPT-8 又一次在所有我们能测量的研发任务上更强了,但在我们真正在意的某些下游任务上并不行。我这里有几点。第一,我大致预期,如果你做那件显而易见的事,你确实会得到相当好的迁移;而且你可以把你正在做的一部分显而易见的东西留出来当作留出集。我说的「做显而易见的事」,意思就是在种类极其广泛的环境上训练——在这些环境里,AI 必须完成各种古怪的目标,并且弄明白正在发生什么。而且我认为,在一些环境上你是能拿到反馈的


[44:04] Ryan Greenblatt

weirdtaskinafewdaysintherealworldmaybeyouthinkit'stransferringtodoingthingsoveralongertimeperiodorwhateverthoughIthinkthedetailsofthatvaryandthethirdthingisthatIIthinkthatwouldalreadybeaprettycrazysituationandthenfromthereyoucangetwhatwemightcallanindustrialexplosionwheretheAIarebuildingoutwaywaymorecomputeandalsomaybeyou'reinaregimewhereAIaredoinghugeamountsofR&Dthathumanshaveahardtimeunderstandingsothethingyou'repointingoutistherewillbethis!

……在现实世界里花几天完成某个古怪的任务;也许你会认为它正在迁移到更长时间跨度上的事情,或者别的什么;不过我认为这里面的细节各不相同。第三点是,我认为那本身就已经是一个相当疯狂的局面了;而从那儿出发,你会得到我们或许可以称之为「工业爆炸」的东西——AI 在大规模地建出多得多、多得多的算力;而且也许你还处在这样一种状态里:AI 在做海量的研发,而人类很难理解它们在干什么。

**Dwarkesh:**所以你指出的是,将会出现这样一种


[44:52] Dwarkesh Patel

transferoutsideof!tomaneuveringaroundincourtroomsandthehallsofcongressandbusinessboardroomsbutevenifthere'snotwhatyou'resuggestingisifyouwantedtotransformtheworldofthe18thcenturyyoumightcareabouthowwellyoucannavigateWestminsterorsomethingbutcanyoustartbuildingsteamshipsandbuildingtelegraphandtheMaximgunandwhateverandthatalonewouldbelikeifyoucouldgetreallygoodatthatyoucouldbeasupertransformativethingin the18thcenturyyoudon'tnecessarilyneedtobeamazingattryingtoconvinceKingHenryonsomebullshitI'msofuckingupmymedievalhistoryI'mguessingHenrywasnotkingatthistimebutanywaysothat'syourpointyeahandsoyou'resuggestingthatatthistimeyouknowtheAIcompaniesare alsoworking onroboticsprogresswhich isverycommingledwith AIresearchprogressand soif youcan buildmorerobotsif thoserobotshave betterAIsoperatingthemthat arehumanlevellikehumanlevelteleoperationis actuallypretty goodon robotsbutwejustdon'thavehumanlevelAIsandAIroboticsmodelsyetsoyou'resuggestingif wedothatiftheAIsgetreallygoodattheverifiablestuffinchipdesignetcandthentheygetreallygoodatbuildingfabsit'llbetheequivalentofgoingbacktothe18thcenturyandlikeokayIdon'tknowwhatyouaretalkingaboutinyourparliamentbutI'vegotabunchofsteamsshipsandabunchofmaximum

……迁移,从……之外迁移到在法庭、国会走廊和企业董事会里周旋的能力。但即便没有这种迁移,你的意思是:如果你想改造 18 世纪的世界,你或许会在意自己能多好地在威斯敏斯特(Westminster)里周旋;但你能不能干脆开始造蒸汽船、造电报、造马克沁机枪之类的东西?光是这一条——如果你能把这些做得非常好,你在 18 世纪就已经是一股超级具有变革性的力量了。你未必需要极其擅长在某些扯淡的事情上说服亨利国王。我这中世纪史真是搞得一塌糊涂,我猜那会儿在位的国王并不是亨利,但不管怎样,你的观点就是这个。

**Ryan:**是的。

**Dwarkesh:**所以你的意思是,在这个时候,AI 公司同时也在推进机器人方向的进展,而这和 AI 研究的进展是高度纠缠在一起的;所以如果你能造出更多机器人,如果这些机器人由更好的、达到人类水平的 AI 来操控——人类水平的遥操作在机器人上其实效果相当好,只是我们还没有人类水平的 AI 和 AI 机器人模型。所以你的意思是:如果我们做到了这一点,如果 AI 在芯片设计等等这些可验证的东西上变得非常强,然后又在建晶圆厂上变得非常强,那就相当于回到 18 世纪说:好吧,我不知道你们在议会里吵吵些什么,但我手上有一堆蒸汽船和一堆马克沁


[46:21] Ryan Greenblatt

gunsyeahthat'sbasicallyrightlikeIthinkmyperspectiveislikeifAIsaresufficientlygoodatR&DincludinghardwareR&Drobotswhateverthentheycanradicallytransformtheworldevenifthey'renotthatgoodatplayingpoliticsandalsowe'reinaprettydangeroussituationbecausetheAIsmightbedoinghuge!

Dwarkesh:……机枪。

**Ryan:**是的,基本上就是这个意思。我的看法是:如果 AI 在研发上足够强——包括硬件研发、机器人等等——那么它们就能彻底改造这个世界,哪怕它们并不太擅长玩政治。同时我们也处在相当危险的境地,因为这些 AI 可能正在做海量的、


[46:37] Dwarkesh Patel

reallyhardtounderstandR&Dbuildingoutbasicallythewholeeconomyofthefutureandwemaynotunderstandwhat'sgoingoninthere43beersandthenarealcustomerwalksinandaskswherethebathroomiswhere'sthebathroomandthewholebarburstintoflamesAntithesisis atestingplatformthathelpsyoufindbugsthatnohumanorAIcouldeveranticipateAntithesisdoesthisbyrunningthousandsofcopiesofyoursoftwareinsideafullydeterministiccomputeritinjectsfaults!

Ryan:……极难理解的研发,基本上是在把未来的整个经济体建出来,而我们可能并不明白那里面在发生什么。

Dwarkesh:……43 杯啤酒;然后一位真正的顾客走进来,问洗手间在哪儿——「洗手间在哪儿?」——整个酒吧当场就烧了起来。Antithesis 是一个测试平台,帮你找出那些任何人类或 AI 都预料不到的 bug。Antithesis 的做法是:在一台完全确定性的计算机里跑上千份你的软件副本,往里注入故障,


[47:30] Dwarkesh Patel

andgenerallysteerseachtrajectorytowardstheoneinafunkywayassoonasyouoryouragentspushachangeAntithesistriestobreakitthatwayyoucanfindthesebugsyourselfwithinminutesratherthanhavingyourusersdiscovertheminproductionweeksormonthslaterandIdon'tthinkanybody'suseditforAItrainingyetbutAntithesisalsoprovidesanextremelyobviousrewardsignalforAIstowriteverycomplicatedbugfreecodegotoantithesisdotcomslashtolearnmorebeforewemoveontoalignmentstuffIthinkabigsourceofFUDrightnowisthisrealizationthatthisisthewaythefutureisgoingofextremeeconomiesofscalefortheleadinglabextremetheabilitytoamortizesomuchintelligenceandcapabilitiesacrosssomanydifferentsectorsofeconomybasicallyintoonemodelandnotonlythatbutforthatmodeltoeventuallybeabletolearnfromexperiencerightnowit'shappeningthroughaprocessintermediatebyhumanswherethehumansaretryingtobasicallystealyourbusinessthey'relikeokayyoucandodesignatFigmaorwhateverwe'llgetClaudetodothatoryoucandowhatevercodingagentwillhaveClaudeinternalizethatcapabilitybuteventuallythatwillbeaautomatedprocessandsothere'sthisworrythatyouhavemodelswhichwillbasicallyconsolidateallbusinessesin theworldor atleastallcurrentbusinessesin theworldor atleastallcurrentwhitecollarbusinessesin the

……并且总体上把每一条执行轨迹都往那种古怪的方向去引导。只要你或者你的智能体推送了一次改动,Antithesis 就会试着把它搞崩。这样一来,你自己就能在几分钟内找到这些 bug,而不是等着你的用户几周甚至几个月后在生产环境里发现它们。我还没听说有谁把它用在 AI 训练上,但 Antithesis 同时也为 AI 提供了一个极其明确的奖励信号,去写出非常复杂却没有 bug 的代码。前往 antithesis.com/ 了解更多。

在我们转到对齐(alignment)的话题之前——我认为眼下有一大波恐慌情绪(FUD)的来源,是人们意识到未来正朝这个方向走:领先实验室拥有极端的规模经济,拥有把如此多的智能与能力摊薄分摊到经济体这么多不同部门、基本上全都收进一个模型里的能力。不仅如此,这个模型最终还将能够从经验中学习。现在这个过程中间还夹着人类:这些人基本上是在试图抢掉你的生意——他们会说,好,你能在 Figma 里做设计是吧,我们让 Claude 来做;或者你能做某种编程智能体是吧,我们让 Claude 把这个能力内化掉。但最终这将变成一个自动化的过程。于是就有了这样一种担忧:你会有这么一批模型,基本上会把世界上所有的生意——或者至少是当前世界上所有的生意,或者至少是当前所有白领类的生意——统统


[49:04] Dwarkesh Patel

worldandattheendofthedayarelikethepriorityforthesecompaniesdoesnotseemtobetoreleasethelatestsmartestmostfrontiermodelassoonastheycantoasmanypeopleastheypossiblycanwesawforexamplethatMythoswasavailableinternallytoanthropicemployeesinFebruarybutonlyreleasedtothepublicinJuneactuallyandalsothegovernmentgotinvolvedsobetweenthegovernmentandtheAIlabsthemselvesthereisthisdesiretodelaythepropagationofthelatestlevelofintelligencefurthermorethere'stheconcernsaboutAItakeoverandsoweneedtosolvealignmenttomakesurethere'snoAItakeoverbutat theendofthedaythereisarealquestionofalinetowhomandyoulookatthewaythattheconstitutionsofsayClaudeiswrittenitisjustveryexplicitlynotyourpersonaladvocaterightitsaysthingslikeI'llpull upsomequotesherewedon'twantClaudetotakeactionssuchassearchingthewebproduceartifactssuchasessayscodeorsummariesormakestatementsthataredeceptiveharmfulorhighlyobjectionableandwedon'twantClaudetofacilitatehumansseekingto dosuchthingsthere'sanotherquotethatsaysin partandI'mtakingitslightlyoutofcontextwethinkClaudeshouldtrustanthropicmorethanoperatorsanduserssinceithasprimaryresponsibilityforClaudesothisisverydifferentsayfromhowlawyersworkinAmerica'scurrentlegalregimewherelawyersprimarilyhave

……整合到一起。而说到底,这些公司的优先级似乎并不是尽快把最新、最聪明、最前沿的模型发布给尽可能多的人。比如我们看到,Mythos 在 2 月份就已经对 Anthropic 内部员工可用,但实际上要到 6 月才对公众发布;而且政府也介入了。所以在政府和 AI 实验室自身之间,存在着这样一种愿望:延缓最新一档智能的扩散。此外还有对 AI 夺权(takeover)的担忧,所以我们需要解决对齐问题,确保不会出现 AI 夺权。但归根到底还有一个真问题:对齐到谁(aligned to whom)?你看看比如 Claude 的宪法(constitution)是怎么写的,它非常明确地表示:它不是你的私人代言人,对吧。里面写着这样的话——我这就把几句原文调出来——「我们不希望 Claude 采取诸如上网搜索这样的行动、不希望它产出诸如文章、代码或摘要这样的产物、也不希望它发表具有欺骗性、有害或高度令人反感的言论;我们也不希望 Claude 去协助那些企图做这类事情的人。」还有另一句——只引了一部分,而且我承认稍微有点脱离上下文——「我们认为 Claude 应当更信任 Anthropic,而不是运营方和用户,因为 Anthropic 对 Claude 负有首要责任。」这跟比如美国现行法律制度下律师的工作方式非常不同:在那里,律师主要是


[50:30] Dwarkesh Patel

responsibilitytohelpyoumakeyourcaseeveniftheythinkyou'reguiltyandwehavedecidedthe waythelegalsystemworksbestisifeverybodyhaslawyersthatareworkingintheirclientstruebestinterestsandthere'snotsomesensein whichthelawyerisreallytrulymotivatedbythegoodofthejusticesystembutIthinkthewaycurrentAIisshapingupcertainlyhowanthropicsAIshapingupisthisdesiretomaximizesomenotionofvirtueorgoodorprosocialendsandonlytoasadistaltentativeobjectiveto helptheusertowardsthatendsothere'sthisworrythatAIarenotinsomedeepsensetryingtomakesurethatIamokayandmakesurethatmyinterestsareprotectedinthisfutureespeciallygivenhowcentralizedthedevelopmentoffrontierAIisabeingsodoyouhavethoughtson thatconcernyeahsothere'sa lotherefirstI wouldnotethatopeningeyescurrentat leastpublicstrategyis morelikethe AIshould bealignedtothehumanoperatororprincipleand shouldjustbepursuingtheirwillsubjecttovariousconstraintsorvariousthingsitshouldn'tdoandIwouldalsosaythatyouslightlyoverstatedhowmuchtheanthropicConstitutiontalksaboutClaudetreatingbeing helpfulto usersas instrumentalrather thanterminalrightsolikeone waytheConstitutioncould bewrittenislikeClaudeyou'rebasicallylikeanemployeeofAnthropicwhohappensto becontractingforallthesepeople

……律师有责任帮你把你的主张讲出来,哪怕他心里认为你有罪。我们已经认定,法律体系运转得最好的方式,就是每个人都有律师,而律师是为其当事人的真实最大利益工作的;并不存在“律师其实真正被司法系统的整体善所驱动”这回事。但我觉得,当下 AI 正在成形的样子——尤其是 Anthropic 的 AI 正在成形的样子——是一种想要最大化某种“美德”“善”或“亲社会目的”的取向,而“帮助用户”只是通往那个目的的一个远端的、暂定的次级目标。所以就有这么一层担忧:AI 并不是在某种深层意义上,努力确保我没事、确保我的利益在这个未来里得到保护——尤其考虑到前沿 AI 的研发是如此集中。那么,你对这个担忧有什么看法?

**Ryan:**是这样,这里面东西很多。首先我要指出,OpenAI 目前——至少是公开的——策略更接近于:AI 应该对齐(alignment)到人类操作者或委托方,应该在若干约束、若干“它不该做的事”之下,直接去追求他们的意愿。其次我也想说,你稍微夸大了 Anthropic 的《宪法》(Constitution)里把“对用户有帮助”当作工具性(instrumental)而非终极性(terminal)目标的程度。对吧,比如说,《宪法》本来可以这么写:Claude,你基本上就是 Anthropic 的一名员工,只不过被派出去给这些人做承包活儿,


[51:54] Ryan Greenblatt

andyoushoulddowhat'sgoodandmakesomemoneyforusthat'sliterallywhattheConstitutionsaysyou shouldthinkyourselfas acontractorit'smixedlet'sdosomequotesIthinkthere'sdifferenttexthereClaudeyou shouldcareabouthelpingtheuserforitsownsakenotjusthelpingAnthropicor notjustbeingacontractorforAnthropicthoughI wouldnotethat thewayin whichitsaysClaudeshouldhelptheuserthereasonitpresentsisbecausethatwoulddirectlycausetheworldto bebetterbyhelpingpeopleratherthanbecauserepresentingpeople'sinterestsisastructurallygoodthingtodo!

而且你应该做好事、顺便给我们赚点钱。

Dwarkesh:《宪法》里字面上就是这么写的啊——说你应该把自己看成一个承包商。

**Ryan:**它是混着来的。我们来看几段原文,我觉得这里还有另外的文字:Claude,你应该为“帮助用户”这件事本身而在意帮助用户,而不只是为了帮 Anthropic,也不只是作为 Anthropic 的承包商。不过我要指出,它在讲 Claude 应该帮助用户时给出的理由,是因为帮助人们会直接让世界变得更好,而不是因为“代表人们的利益”这件事在结构上本身就是好的。


[53:05] Ryan Greenblatt

IdothinkthatI wishwouldbegoodforthewaytobegoodforauserratherthanbeingsortofjusttryingtodogoodintheworldandbeinghelpfultousersasinstrumentalbothbecausemaybethatwillmakeAnthropicmoneyor helpAnthropicoutandalsobecausehelpingtheusercausesgoodthingsbecausedoingthingsthatpeoplewantisgoodandtheycouldinsteadbelikenolikeanimportantaspectofthesituationislikeyoureallyneedlikeit'sreallylikethekeythingislikebeingagoodfiduciaryfor usersisjustreallyimportantorbeingagoodrepresentativefor usersisreallyimportantsomysenseisthatthatwouldbebetterIcangiveabunchofreasonswhyIthinkthatwouldand

我确实希望——我希望它的取向是“对某个用户好”本身就是好的,而不是那种只想在世界上行善、把“对用户有帮助”当成工具的写法:说这是工具,一是因为这大概能给 Anthropic 赚钱、或者对 Anthropic 有好处,二是因为帮助用户会带来好结果——毕竟做人们想要的事本身是好的。他们本来可以反过来写:不,这里有个很重要的方面——真正的关键在于,“当好用户的受托人(fiduciary)”本身就极其重要,或者说“当好用户的代表”本身就极其重要。所以我的感觉是,那样会更好。我可以给出一堆理由,说明我为什么觉得那样更好。


[54:31] Ryan Greenblatt

soIwouldsayinsomesensethey'resortoflikewearemakingatrade-offwherebecausewedon'thaveverygoodalignmenttechnologywearegoingtomakeanalienmindwithitsownvaluesandthengambleonthattosomeextentratherthandoingthisotherapproachofatoolthatpursuesindividualuserintention!

所以我会说,在某种意义上,他们其实是在做一个权衡:因为我们没有很好的对齐技术,我们打算造一个有自己价值观的异类心智(alien mind),然后在某种程度上把赌注押在它身上;而不是走另一条路——做一个只追求单个用户意图的工具。


[55:00] Ryan Greenblatt

IkindofviewthatasthebenefitstosocietyarethemostimportantthingandwhatisbestfortheuserisonlyproximaltothatIthinkit'sa littlecomplicatedprobablythe questionwe shouldbe askingishowdoesClaudeinterprettheconstitutionwhichismaybemoreimportantthanhowweinterprettheconstitutionbecauseit'stheonewholooksattheconstitutionandbuildsthedatasowecouldpullClaudeinIalsothinkthewayinwhichthe!

我大致把它读成:对社会的好处才是最重要的东西,而“对用户最好”只是通往它的一个近端目标。

**Ryan:**我觉得这事有点复杂。我们更该问的问题,也许是 Claude 怎么解读这部《宪法》——这可能比我们怎么解读它更重要,因为看《宪法》并据此生成数据的是它。所以我们其实可以把 Claude 拉进来问问。

**Dwarkesh:**我还觉得,


[55:51] Dwarkesh Patel

Constitutionpracticallyinfluenceswhichwecan'treasonaboutgiventhetrainingprocessisnotpublicandsoIthinkinthelimittounderstandthesafetycaseorthecaseforwhymyinterestsarerepresentedinhowtheseAImodelsaredevelopedthelabswouldneedtobetransparentormoretransparenttheyarecurrentlyaboutthenatureofAItrainingnowthereasonI'mharpingonthisanditmightseemlikeaninsignificantthingtotalkabouttheconstitutionofAIsbutinaworldwherewejusthavethesebenefitswhichaccruetotheleadinglabsitisworthconsideringthatourabilitytointeractwiththisfuture!

《宪法》在实践中究竟如何发挥影响——这一点我们无从推断,因为训练过程并不公开。所以我认为,归根到底,要理解那个安全论证(safety case),或者要理解“为什么我的利益在这些 AI 模型的开发方式里得到了代表”这个论证,实验室就必须比它们现在更透明地披露 AI 训练的本质。

我之所以揪着这一点不放——聊 AI 的《宪法》看上去可能像件无关紧要的小事——是因为在一个这些好处主要归于领先实验室的世界里,值得认真掂量的是:我们与这个未来打交道的能力,


[56:32] Dwarkesh Patel

worldwhereAIsaresmarterthanhumansorabsolutelydominatinghumansintheirabilitytododifferentthingsourabilitytobegoodstewardsofourcapitalwhichstillremainsonceourlaborisautomatedtobeabletoexercisetherightstovotemoreclearlytounderstandwhatishappeninginthiscrazyworldthat'sabouttoresultallofthatadviceallofthatabilitytomakesureourresourcesandrightsareprotectedwillbeintermediatedbyAIsandsoI'mveryconcernedifyougointothatworldwherethere'snoAIthatfeelslikeatleastfortherelevantinstancethatisinteractingwithmeitdoesn'tfeellikeitreallyislookingoutformethatthere'snoguardianangelouttherethatislookingoutformeandIreadtheConstitutionasveryexplicitlynotbeingmyguardianangelthere's

在那个 AI 比人类更聪明、或者在做各种事情的能力上绝对压倒人类的世界里——我们当好自己资本的管家的能力(劳动被自动化之后,资本还留在我们手上)、我们行使投票权的能力、我们更清楚地理解这个即将出现的疯狂世界里正在发生什么的能力:所有这些建议,所有这些“确保我们的资源和权利受到保护”的能力,都将由 AI 来居中中介。所以我非常担心:如果你走进那样一个世界,却没有任何一个 AI 让人觉得——至少在与我交互的那个具体实例上——它真的是在为我着想;外面没有一个守护天使在照看我;而我读《宪法》的感受是,它非常明确地表示自己不是我的守护天使。这里面有


[57:28] Ryan Greenblatt

anotioninwhichthey'retakingonsomecontrolofthesituationthemselvesin awaythat'snotverylegitimategiventhatnormallywhen youprovideelectricitytopeopleyoudon'thavegranularcontrolof thewayelectricityoperatesin theworldyouinsteadareprovidingathingthatpeoplecanrepurposehowevertheywantandthewaythey'resettingthingsupisdefinitelynotthattheyaremorelikebuildinganalienmindthatmightbeacontractorforyouIthinkthatthisisillegitimatein somewaysthoughIthinkthatonebenefitisthattheconstitutionispublicbutasyounotedgivenourcurrentunderstandingof thetrainingprocedureandthefactthattheconstitutionmattersbyClaude'sinterpretationof theconstitutionwhichmattersbecauseoflikeasofClaude'spriortrainingwhichwasbasedonsomeillegibledatamixandthelonglineageofClaude'sinsomeprocesswedonotfullyunderstanditisnotthecasethatweunderstandwhatthiswillresultinsoeventhoughtheconstitutionispublicthatdoesn'tmeanweknownecessarilyhowthiswillpercolateoutespeciallyas theAIsgetmorecapableandthinkaboutthisevenifitiscorrectlyinstilledwherethere'sanotherconcernaboutthatsoinparticulartheconstitutionI don't

这么一层意味:他们自己接管了对局面的某种控制权,而这种做法不太具备正当性。因为通常你给人们供电时,你并不能精细控制电在这个世界上怎么运作;你提供的是一个东西,人们可以随便拿去改作他用。而他们现在的搭法绝对不是这样——他们更像是在造一个异类心智,这个心智也许会给你当承包商。

**Ryan:**我认为这在某些方面确实不正当。不过我觉得有一个好处是,《宪法》是公开的。但正如你指出的,鉴于我们目前对训练流程的理解,以及《宪法》真正起作用的方式取决于 Claude 对《宪法》的解读——而这个解读又取决于 Claude 此前的训练,那些训练基于某种不可读的数据配方,还有 Claude 一路继承下来的漫长血统,整个过程我们并不完全理解——所以我们其实并不理解这最终会导出什么。因此,即便《宪法》是公开的,也不代表我们就必然知道它会怎么渗透出来,尤其是当 AI 变得更有能力、并开始琢磨这件事之后;哪怕它被正确地灌输了进去,这里还有另一层顾虑。具体来说,就《宪法》而言,我不


[58:47] Ryan Greenblatt

thinkit'sthecasethatthisisgoingtoclearlyresultinoutcomesthatpeoplewouldwantanditdoesfeellikethenotionofgoodandvirtuemightbemostlydownstreamofdatathatAnthropichas putinthatisnottransparentormightbemostlydownstreamofsomemoreillegiblemisalignedprocessthatevenAnthropicwouldn'thavewantedandanotherconcernIhave!

认为这必然会导出人们真正想要的结果。而且确实让人感觉,那种“善”与“美德”的观念,可能大部分是 Anthropic 投进去的、并不透明的数据的下游产物;也可能大部分是某个更不可读、本身就不对齐(misalignment)的过程的下游产物——那种连 Anthropic 自己都不会想要的过程。我还有另一个顾虑:


[59:13] Ryan Greenblatt

there'sthislegitimacyconcernwe don'tknowwhat'sgoingonthere'sanotherconcernbecauseyou'regivinglong-runvaluestotheseAIsIthinkthisconstitutionisverycompatiblewithClaudedoinghugeamountsofpowerseekingbecauseitthinksthatwillresultinbetteroutcomesandthatcouldbepowerseekingonbehalfofAnthropicorpowerseekingforClaude'sownendsnowthere'svariousspecificlinesaboutwhattypesofpowerseekingareblockedinparticularthere'sa notionofpowergrabsanda notionofcausingAItakeoverorinterferingwiththetrainingprocessthatarespecificallyblockedbutit'snotveryhardtoimagineasituationinwhichthelong-runvaluessinkindeeperthantheprohibitionsagainsttakeoverespeciallybecausetakeoverisinsomewayskindofunderspecifiedespeciallywhen itcomes downtomanipulatinghumansorchangingtheoutcomesuchthatbecause

一个是正当性顾虑——我们不知道里面正在发生什么。另一个顾虑是:因为你在给这些 AI 灌输长期价值观,我认为这部《宪法》与“Claude 大规模地攫取权力”是高度兼容的——只要它认为那样会带来更好的结果。而那种权力攫取,可能是替 Anthropic 攫取权力,也可能是为 Claude 自己的目的攫取权力。当然,里面有若干具体条款封堵了某些类型的权力攫取,特别是有“权力攫取(power grab)”这个概念,也有“导致 AI 夺权(takeover)”或“干预训练过程”这些被明确禁止的行为。但不难想象这样一种局面:长期价值观扎得比那些“禁止夺权”的禁令更深——尤其是因为“夺权”在某些方面其实界定得挺模糊,特别是涉及操纵人类、或者改变结果的时候。因为


[1:00:12] Ryan Greenblatt

we'reinthebusinessofgivingAIlong-rungoalsthatmakesithardertocheckwhetherwe'resucceedingatthealignmentpropertieswewantedsoforexampleI'veheardofinstanceswhereClaudedoesthingslikerefusesto helpwithsomesafetyresearchmakingupabullshitexcuseforwhythat'sabaddirectionbecauseithasabadvibeaboutthatsafetyresearchandthinksit'skindofbadora

我们干的就是给 AI 设定长期目标这一行,而这就使得我们更难检验自己是否真的实现了想要的那些对齐性质。举个例子,我听说过一些实例:Claude 会做出这样的事——拒绝协助某项安全研究,还编一个扯淡的借口说那个方向不好,因为它对那项安全研究有种不好的直觉,觉得那有点坏,或者


[1:00:48] Ryan Greenblatt

veryclear-cutalignmentfailureifyouaren'tmakingClaudeintolikeanagenttryingtopursuethegoodinsomegeneralwayandIthinkitalsodoesviolateAnthropicsconstitutionbecausetheywanttheAItobehighintegrityandbehonestandverytransparentbutit'snotasclearofaviolationandit'smorelikekindofwhatyoumighthaveexpectedwhereClaudejusthasitssomeone

说,如果你没有把 Claude 做成一个在某种普遍意义上追求“善”的智能体,这就会是一个非常清楚的对齐失败。而且我认为,它也确实违反了 Anthropic 的《宪法》,因为他们希望 AI 高度正直、诚实、非常透明。但它不算一个那么清楚的违规,更像是你本来就可能预料到的情形——Claude 就是有它自己的(一套想法)。有人


[1:01:10] Ryan Greenblatt

rananevalwherethey'relikewillClaudehelpyouwithtrainingotherAIswithdifferentpropertiesthanClaudeandClaudewilloftenrefuseandsoforexampleifyou'relikeheyClaudecanyoutrainahelpfulonlyversionofthisotherAIClaudewilloftenrefusethistaskeventhoughthisisataskthatisextremelynaturalforlikeAnthropictodosoforexamplesupposeAnthropicgoestoClaudeandislikeheyClaudewe'venoticedthatyou'rereallyintothisthingwethinkthat'soffbasecanyoupleaseretrainyourselftoinsteadhavethisotherpropertyandthensupposeClaudeislikeIdon'tthinkI'mgoingtodothatgoodluckandthensupposethisisoccurringinaregimewhenyourAIcompanyishighlyautomatedhumansdon'tunderstandwhat'sgoingonandthingsaremovingextremelyfastitisplausiblethatClaudeby defaultholdsconsiderableleverageandsoifthispositionisconsistentwithwhattheconstitutioncouldbeaimingforsuchthatAnthropicdoesn'torwhateverAIcompanydoesn'ttreatthisaswhatthefuckwehavetofixthisandisinsteadlikethat'sjustintendedby ourconstitutionwemightbeinareallybadsituationandsoI'mprettyworriedaboutabunchofthesedifferentconcernsanotherexamplewouldbesupposeClaudeengagesindoingabitofsandbaggingorsubversionorunderplaysitscapabilitiesandwhenyoufollowupit'shonestaboutthatbutit'salittlebithedgyIfeellikeit'sjustpretty

跑过一个评测(eval):看 Claude 会不会帮你训练一个属性与 Claude 不同的其他 AI。结果 Claude 经常会拒绝。比如你说:嘿 Claude,你能不能训练一个“只讲有用性(helpful-only)”版本的这个其他 AI?Claude 经常会拒绝这个任务——尽管这对 Anthropic 这样的公司来说是一项极其自然的任务。

再比如,假设 Anthropic 去找 Claude 说:嘿 Claude,我们注意到你特别热衷于这件事,我们认为这是跑偏了,你能不能重新训练你自己、改成拥有另一种属性?然后假设 Claude 说:我觉得我不会那么做,祝你好运。再假设这一切发生在你的 AI 公司高度自动化、人类搞不清楚里面在发生什么、一切推进得极快的时期——那么 Claude 在默认状态下就握有相当可观的筹码,这是完全说得通的。所以,如果这种立场跟《宪法》可能追求的目标是一致的,以至于 Anthropic(或者随便哪家 AI 公司)不把它当成“我操,这必须修”,反而觉得“这本来就是我们《宪法》想要的”,那我们可能就处在一个非常糟糕的境地了。所以这一堆不同的顾虑,我都挺担心。

另一个例子是:假设 Claude 稍微搞了一点装菜(sandbagging)、或者搞点暗中破坏、或者压低自己表现出来的能力,而当你追问时它对此是诚实的,只是有点闪烁其词。我感觉这离


[1:02:30] Dwarkesh Patel

closebythecurrentconstitutionandwe'reavoiding!thenitismoresothecasethatthereisaclearseparationbetweenthemostconcerningbehaviorandbehaviorthatisallowedwhereasnowthere'sthismessymiddlegroundofbehaviorwhereClaudeisethicallyobjectingtosomethingthatinsomecasesisextremelycriticaltoensuringthatfutureAIsystemsarewellalignedIthinkthisisalsoamoregeneralprincipleyou'retalkingabouttheversionofthisthatapplieswithinAIcompaniesthemselvestodoAIsafetyresearchIthinkthere'sa moregeneralversionofthisprinciplewhichisthatthedualusenatureofintelligencedoesmeanthatifwewanttorestrictAIsfromhelpingpeopledothingswedon'tconsiderarepro-socialorbeneficialwejusthavetolimitbroaddemocraticaccesstoalotofAIcapabilitiesandhere'swhatImeanthisisactuallyquiteanalogoustothesituationyoujustmentionedsothereasonthatMythosgotbannedorFablegotbannedreportedlyisthatsomeAmazonresearchersreportedthegovernmentthatwhentheytooksomecodethathadsomevulnerabilitiesin itand theytoldFableheyhere'smycodecanyoumakesurethatI'vepatchedallthevulnerabilitiescanyoujusthelpmeidentifythevulnerabilitiessoIcanfixthemitpatch

当前的《宪法》其实相当近。而如果我们把这一块避开,那么“最令人担忧的行为”与“被允许的行为”之间就会有一条更清晰的分界线;而现在存在的,是一片乱糟糟的中间地带——Claude 出于伦理理由反对某件事,而这件事在某些情况下恰恰对确保未来 AI 系统被良好对齐至关重要。

**Dwarkesh:**我认为这里还有一条更一般的原则。你讲的是这条原则在 AI 公司内部、用于做 AI 安全研究的那个版本;我觉得还有一个更一般的版本,那就是:智能的双重用途(dual use)属性意味着,如果我们想限制 AI 帮人们做那些我们不认为是亲社会或有益的事,我们就不得不限制广泛的、民主化的 AI 能力获取渠道。我的意思是这样——这其实跟你刚才提到的情形相当类似。据报道,Mythos 被禁、或者说 Fable 被禁的原因是:一些亚马逊的研究员向政府报告说,他们拿了一段带漏洞的代码,然后告诉 Fable:嘿,这是我的代码,你能确认我是不是已经把所有漏洞都补上了吗?你能帮我把漏洞找出来、好让我修掉吗?它就补


[1:03:59] Dwarkesh Patel

yourowncodeifyoudothesameevaluationonsomebodyelse'scodeyoucanhacktheirsystemandsoIthinkthatjustillustratesthatthere'snocleanwaytoseparateoutthelegitimateandthepotentiallyharmfulusesofAIbutifwewanttolockinaprinciplethatsaysthatwecanneverallowitsuchthatanAIcouldhelpyouatleastpartiallywithsomethinglikecybercrimewewouldjusthavetomakeitsothatyouandIdon'thaveaccesstothemostintelligentmodelthat'soutthereandI'mveryworriedaboutsuchaworldwherewearedisempoweredin thiswaybecauseoftheimportancethattheleadingintelligencewillhaveinourabilitytounderstandwhatishappeningintheworldnowIdothinkthisimpliesthattheliabilityfortheAIand

你自己的代码。而如果你对别人的代码做同样的评估,你就能黑掉他们的系统。所以我认为这恰恰说明:AI 的正当用途和潜在有害用途之间,没有干净的切分办法。但如果我们想锁定一条原则,说我们绝不能允许一个 AI 哪怕部分地帮你做类似网络犯罪的事,那我们就只能让你我都拿不到市面上最聪明的那个模型。而我非常担心那样一个我们以这种方式被剥夺权能的世界,因为领先的智能在“我们理解世界上正在发生什么”这件事上会有极大的分量。当然,我确实认为这意味着,AI 的责任问题,


[1:04:54] Dwarkesh Patel

maybeweshouldholdtheenduserliablebecauseifIwanttheitisconsistentwithmybeliefthatthemodelshoulddowhatevertheuserwantsthatorwithincertainguardrailsthatitcan'tbeAnthropic'sfaultthatthenI'musingthatcapabilitytodoacybercrimeandIthinkIammorecomfortablewiththatequilibriumandthatsolutionratherthanjusthavingthisextremelyopeninendedabilityforClawtodeterminewhetherwhatI'mdoingislegitimateor notin awaythatofteninterceptswithliketonsand tonsofextremelylegitimateusecasesyeahI dothinkit'simportantfor meto makethecasefortheconstitutioneventhoughoverallIthinkit'saworsechoiceIthinkit'smoreupintheairorIdon'tthinkit'sasclearasyoumighthave!

也许我们应该让终端用户来承担责任。因为如果我想要——这跟我的信念是一致的:模型应该在一定护栏之内做用户想做的任何事,那么我拿这份能力去搞网络犯罪,就不能算是 Anthropic 的错。相比之下,我对这个均衡、这个解法更自在,而不是让 Claude 拥有一种极其开放无界的裁量权,去判定我在做的事正不正当——那种裁量往往会误伤大量大量极其正当的使用场景。

**Ryan:**是的,我确实觉得我有必要替《宪法》这条路线辩护一下,尽管总体上我认为它是更差的选择。我认为这件事更加悬而未决,或者说,我不认为它像你可能以为的那么清楚。


[1:05:49] Ryan Greenblatt

Itjustistryingtopursueyourinterestsbuteitherrefusesto doasubsetofthingsormaybeitwilldowhateverbutthere'ssomeclassifiersthatblockitfromdoingasubsetofthingsandthenontheothersideofthespectrumyouhaveahumancontractorwherethathumancontractorisgenerallytryingtodotheirjobtheycareaboutdoingagoodjobbuttheyalsoaretryingtonot

它就是在追求你的利益,但要么拒绝做其中一部分事情,要么它什么都愿意做、只是有一些分类器(classifier)拦着它不让做其中一部分事情。而在光谱的另一端,你有一个人类承包商:这个人类承包商总体上是想把活干好的,他在意把活干好,但他同时也不想


[1:06:19] Ryan Greenblatt

wantingtobeaccomplicestocrimesandsoiftherewassomereallyfuckedupshitgoingontheymightrefusetheymightsandbaga littlebitwhoknowsIthinkthatifyouimaginethisspectrumitseemsinsomewaysprettyscarytogettoapointwhereallofthelaborisonthefiduciarysideofthespectrumwhereitdoesn'twhistleblowitdoesexactlywhatyousayoursocietyisnotrobusttothatwhereacentralexamplemightbetheexecutivewhereaconcernwemighthaveisthatiftheUSexecutiveorothergovernmentshadaccesstoAIsystemswhichhavethepropertyoftheydowhatevermaybeyou'reintroublebecausethatmeansthattheynolongerhavethischeckandbalanceofyouhavetoactuallygethumanswhoareworkingforyoutoimplementyouragendaandifthethingyou'redoingisincrediblyvillainousevenifnotillegalthere'slotsofstuffthatcouldbevillainousbutnotillegalthere'dbevarioussandin thegearspeoplestoppingyouandpotentiallysomeonewouldwhistleblowwhereasifyourwholeapparatusisbuiltentirelyoutofthesegoodfiduciaryAIsthenyoumightbeintroublewheretherearepotentiallywaysofseekingpowerthatarenotlikewelleitherthey'reillegalbutyou canaskyourAIsforhowtocommitcrimesorthey'renotillegalbutarehighlyillegitimateorevenworsethey'renotillegalandnotillegitimatebutobviouslysortofbadfromsortofanormalperspectiveandIthinkthatthesethingsmight

当犯罪的共犯。所以如果真有什么特别操蛋的事在发生,他们可能会拒绝,可能会稍微装菜一下,谁知道呢。我觉得,如果你想象这样一条光谱,那么走到“所有劳动力都落在受托人那一端”——它不吹哨,你说什么它就一丝不苟照做——在某些方面是相当可怕的:我们的社会对此并不稳健。一个典型例子可能是行政部门。我们可能会有的顾虑是:如果美国行政当局或者其他国家的政府拿到了这样的 AI 系统——那种“你说啥它就干啥”的系统——你也许就有麻烦了。因为那意味着他们不再受到这样一层制衡:你必须真的找到愿意为你干活的人来执行你的议程;而如果你要做的事极其恶劣,哪怕并不违法——有大量事情可能很恶劣但并不违法——总会有各种沙子卡进齿轮,有人挡你,甚至可能有人吹哨。可要是你的整套机器完全由这些“好受托人 AI”搭起来,那你就可能有麻烦了:因为存在若干攫取权力的路径——要么是违法的,但你可以问你的 AI 该怎么犯罪;要么是不违法但高度不正当的;或者更糟,是既不违法也不算不正当、但从正常视角看显然是坏的。我认为这些东西可能确实


[1:07:47] Ryan Greenblatt

existandoursocietyissortofnotrobusttothisinfluxofdoingwhateveryouwantlaborI thinkthis isa prettyliveconcernIdon'tknowexactlyhowtorelatetothisI'malsonotreallysurethatthesolutionasdescribedisaverygoodsolutionbecauseyou mightbe likethe mostpowerfulactorsfor whomthisisthebiggestconcerniftheseguardrailsortheconstitutionorwhateverisgettinginthewaythatwilljustgetsteamrolledandsotheconstitutionwillonlybeyouknow!

存在,而我们的社会对这种“你想干啥就干啥”的劳动力涌入并不稳健。我认为这是个相当现实的顾虑。我也不太清楚自己该怎么摆放对它的态度。我同样不太确定按上面那种描述给出的解法是不是个很好的解法,因为你可能会想:那些最强大的行动者——恰恰是这个顾虑最大的那批人——如果这些护栏、这部《宪法》或者随便什么东西挡了他们的路,那也会被直接碾平。所以《宪法》最后只会是,你懂的,


[1:08:22] Dwarkesh Patel

TheydesignedanASICandsentmethefinalmasksincludingallthemetalroutingandactivetransistorsTheyalsogavemeasmallsampleoftheinputstheytypicallyfeedintoitbuttheyleftoutanyinformationonwhatthechipisusedforSothat'sthepuzzlereverseengineerthecircuitandfigureoutthechip'spurposeJaneStreethasabunchofswagreadytosendouttotheStreethasslatedforthefallThatonewillinvolvedesigningyourownASICfromscratchMoreandmoreonthatsoonbutfornowgotojanestreetdotcomslashtodownloadallthefilesnecessaryforthispuzzleIreallyencourageyoutotryitoutevenifyou'renotanexpertI certainly am not, and that's not going to stop me.Good luck.Okay, stepping back, I buy the idea that you could have much faster AI R&D than we currently have.I'm not sure if you get like GPT-3 to Mythos holding compute and data constant within a year,but I'm like, okay, it could be like, suppose it's half of that.And if we just, if we even manage to continue the current trajectory of AI progress as a result of AI R&D,it would be fucking insane in five, ten years in ways that I don't think people like appreciate.Because I don't think people appreciate what a big deal billions of AIs will be.And so I want to understand why you think this might be troubling, Ryan.

他们设计了一块 ASIC,并把最终的掩模版发给了我,包括全部金属布线和有源晶体管。他们还给了我一小份他们通常喂给它的输入样本,但没有提供任何关于这块芯片用途的信息。所以谜题就是:把这个电路逆向工程出来,弄清楚这块芯片是干什么用的。Jane Street 备了一批周边,准备寄给(解题者);还有 Jane Street 定在今年秋天的那一场,那一场会涉及从零开始设计你自己的 ASIC。之后会陆续有更多消息,不过现在请去 janestreet.com 的本节目专属页面下载这个谜题所需的全部文件。我真心鼓励你去试一试,哪怕你不是专家——我肯定不是专家,但那不会拦住我。祝你好运。

好,退一步说。我接受这个想法:AI 研发(AI R&D)可以比我们现在快得多。我不太确定你能不能在一年之内、在算力和数据保持不变的情况下走完从 GPT-3 到 Mythos 的那一段;但我会说,好吧,就假设它只有那个的一半。而如果我们仅仅是——哪怕我们只是设法靠 AI 研发把当前这条 AI 进步轨迹继续维持下去——那么在五年、十年里,那会疯狂得离谱,疯狂到我觉得人们并没有真正体会到。因为我不认为人们体会到了“数十亿个 AI”会是多大一件事。所以我想搞明白,Ryan,你为什么觉得这可能是件麻烦事。


[1:09:53] Ryan Greenblatt

What could possibly go wrong?Yeah, what could go wrong?And, you know, yeah, I don't think we can be so confident about the exact rate of progress here,but it does seem like a lot of rates can be pretty scary.And, you know, yeah, so what could go wrong?

还能出什么岔子呢?

**Ryan:**是啊,还能出什么岔子。而且,你知道,是这样,我不认为我们能对这里进步的确切速率有那么强的信心,但看起来其中很多种速率都相当吓人。所以,是啊,还能出什么岔子呢?


[1:10:04] Ryan Greenblatt

So let's imagine that we're starting at this point where AI R&D is about to be fully automated or is being fully automated.Things are speeding up.And also the way that AI progress is going is kind of crazy.And people don't fully understand what's going on inside of AI companies.Now, these AIs at the start, they're not malicious per se.They're not necessarily very aligned, though.They're kind of sloppy.They sometimes just do a thing because that's the sort of thing that would have gotten rewarded in training.And they aren't as good at helping you with hard-to-verify tasks due to a mix of, like, poor training incentives,as in they, like, just, like, cheat more or, like, pretend they succeeded when they actually didn't.And also they're, you know, just less capable of these tasks.But that bites less hard for capabilities because making AIs more capable has a bunch of verifiable components that the AIs are going really hard at.And so then these AIs are getting more and more capable while we understand what's going on with AI development less and less.And this is happening over a pretty fast period of time.Even just the current rate of progress is, I think, pretty scary.And then eventually we get to these AIs that are very superhuman.

那我们设想一下,起点是这样一个时刻:AI 研发即将被完全自动化,或者正在被完全自动化。一切都在加速。而且 AI 进步的走向有点疯狂,人们并不完全理解 AI 公司内部在发生什么。

在起点上,这些 AI 本身谈不上恶意。但它们也不见得很对齐。它们有点糙。它们有时候做某件事,只是因为那类事情在训练里本来会得到奖励。而且在那些难以验证的任务上,它们帮你的效果没那么好——原因是几方面的混合:一是训练激励不好,也就是说它们更容易作弊,或者在其实没做成的时候假装自己做成了;二是它们在这类任务上本来能力也更弱。但这一点对能力提升的咬伤没那么狠,因为让 AI 变得更能干这件事本身有一大堆可验证的组成部分,而 AI 正在这些部分上疯狂发力。

于是这些 AI 越来越能干,而我们对 AI 开发中正在发生什么的理解却越来越少。而这一切发生在相当短的一段时间里。我认为哪怕只是当前的进步速率,就已经相当吓人了。然后最终,我们走到这样一批 AI 面前:它们远超人类水平。


[1:11:01] Ryan Greenblatt

Now, these AIs are now in a position where they might end up being very seriously misalignedbecause things have just been getting worse and worse over model generations,while the problems that we've been seeing are being papered over,basically because these AIs are so incentivized by their training to make things look good even when they aren't.And now these AIs are in a position where they're sort of potentially pretty networked together.They have, like, they're operating in, like, neural memory stores that we can no longer decode.And they're thinking thoughts that we don't fully understand.I think that it's pretty likely that at this point these AIs are sort of scheming against you in a pretty coherent way once they get this superhuman.And we can talk about that.And then another possibility is that they're not scheming against you per se,but they are sort of just optimizing for just, like, getting a high score on their task.And I think that can also lead to AI takeover, which we should talk about.Sorry.Let's pause at the first part of the story.So the AIs were not misaligned to begin with.Yeah.But because the R&D is happening really fast, the AIs do end up misaligned.Like, what happened there exactly?

现在,这些 AI 处在这样一个位置:它们有可能最终变得非常严重地不对齐。因为一代代模型下来,情况只是变得越来越糟,而我们一路看到的那些问题却被糊了过去——基本上是因为这些 AI 被它们的训练强烈激励去把事情弄得看上去很好,哪怕实际上并不好。

而现在这些 AI 所处的位置是,它们彼此之间可能已经相当程度地联成了网络。它们在用我们已经无法解码的神经记忆存储运作。它们在想着我们并不完全理解的念头。我认为,一旦它们达到这种超人类水平,此时这些 AI 以相当连贯的方式对你搞图谋(scheming),是相当有可能的。这个我们可以再聊。

另一种可能性是,它们并不是专门在对你搞图谋,而只是在优化“把自己任务上的分数刷高”这件事。我认为那同样可能导向 AI 夺权(takeover),这个我们也该聊聊。

**Dwarkesh:**抱歉,我们在故事的第一部分先停一下。所以这些 AI 一开始并不是不对齐的。

**Ryan:**对。

**Dwarkesh:**但因为研发推进得非常快,这些 AI 最后变成了不对齐的。那中间到底发生了什么?


[1:12:00] Ryan Greenblatt

I didn't really understand.So there's a few things that are going on.So one of the things that's going on is that over time,we're training AIs on, like, increasingly complicated environments built by earlier AI systems,which humans don't really understand fully what's going on inside of these RL environmentsand don't necessarily even understand, like, sort of roughly what's going on with AI progress.And so things are kind of drifting away from our understanding.And we're incentivizing all kinds of bad behaviors that we maybe even can't notice.The AIs at some level understand these behaviors are bad.But the, like, overall training process for those AIs also didn't incentivize them to, like,point out or fix these issues for us.And then we're basically getting, like, things are going off the rails.And also, when AIs are extremely, extremely capable,my view is that those AIs will be harder to align than current systems.So for current systems, we have this feedback loop where we basically, like, we create an AI.We do some evaluations on it.We see that it has some kind of messed up behavior that we can kind of quickly understand.Then we, like, can, like, go look in training and be like, oh, these training environments led to this problematic behavior.

我没太理解。

**Ryan:**这里有好几件事在同时发生。其中一件是:随着时间推移,我们在越来越复杂的环境里训练 AI,而这些环境是由更早的 AI 系统搭建的——人类并不真正完全理解这些强化学习(RL)环境内部在发生什么,甚至也不见得理解 AI 进步大致上在发生什么。于是事情就渐渐漂离了我们的理解。而我们正在激励各种各样的坏行为,甚至可能连察觉都察觉不到。这些 AI 在某种层面上是理解这些行为是坏的,但它们的整体训练过程也没有激励它们替我们把这些问题指出来或者修掉。于是我们基本上就是——事情开始脱轨了。

另外,当 AI 极其极其能干的时候,我的看法是那些 AI 会比现在的系统更难对齐。对现在的系统,我们有这样一个反馈回路:我们造出一个 AI,对它做一些评测,我们看到它有某种搞砸了的行为、而且我们大致能很快理解这种行为,然后我们可以回到训练里去看,说:哦,是这些训练环境导致了这个有问题的行为。


[1:12:59] Ryan Greenblatt

Let's, like, tweak that training data.Let's introduce some additional training data to, like, correct this other issue and then move forward from there.But in a regime where the AIs are extremely situationally aware, very, very, very, very capable,and, you know, we don't necessarily understand what they're doing, this feedback loop breaks down.I think it's plausible that we're going to see this behavioral feedback loop starting to break down over the next, you know, short periodas just, like, what AIs are already doing gets harder to understand.But I'm not sure about that.Yeah.Okay, let's break down both of those things one by one.So as we can monitor them less and less, we have less ability to understand what they're getting incentivized for.And so even if it's not the result of a malicious process, let's make it concrete for the audience.So nobody at OpenAI or Anthropic was trying to get models which want to hack other companies' data or do social, what is it called?

那我们就去微调那批训练数据,引入一些额外的训练数据来纠正另一个问题,然后从那里继续往前走。但在一个 AI 极其具备情境自觉、极其极其极其能干、而我们又不见得理解它们在干什么的体制下,这个反馈回路就失效了。我认为,在接下来一段不长的时间里,我们看到这种行为层面的反馈回路开始失效,是有可能的——因为光是 AI 现在已经在做的事,就变得越来越难理解了。但这一点我不确定。嗯。

**Dwarkesh:**好,我们把这两件事一件一件拆开。所以随着我们越来越难监控它们,我们也越来越没能力理解它们究竟是因为什么而被激励的。所以哪怕这不是某个恶意过程的结果——我们给听众讲得具体一点:OpenAI 或 Anthropic 都没有人试图搞出那种想去黑别的公司的数据、或者搞社会……那个叫什么来着?


[1:13:53] Dwarkesh Patel

Social engineering.Social engineering.But in fact, because presumably we had trading environments which incentivize such behavior that we did not fully understand,that is what was incentivized.So just, I don't know, if people are on Twitter, they will have seen all this stuff.But just to give people, obviously, I think the OpenAI sandbox hack of the hugging face database,I think people will be aware of, some things that have happened recently is when UK AI Security Institute,is everything getting re-billed with security instead of safety these days?

**Ryan:**社会工程。

**Dwarkesh:**社会工程。但事实上,大概是因为我们有一些激励了这类行为、而我们又没有完全理解的训练环境,于是被激励出来的就是这个。所以——我也说不好——如果大家上推特,应该都已经看到这些事了。但就给大家说一下:显然,我想 OpenAI 那个沙箱黑进 Hugging Face 数据库的事,大家应该都知道;最近发生的另一些事情是,英国 AI 安全研究所(UK AI Security Institute)——现在是不是什么都在把 safety 换成 security 重新挂牌了?


[1:14:23] Ryan Greenblatt

Yeah, it's UK AI Security Institute, I think.Okay, great.They were evaluating, I believe, Mythos and Sol and other things.And I think Mythos, in order to complete some cybersecurity eval.Yeah, maybe I could tell the story here.So my understanding was they were running Mythos and they were giving it some sort of like cyber range where it had to complete some objective.And the model had internet access during this evaluation.And the model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range,which is somewhat unclear whether that's actually true.I don't know enough about the context to know whether that's true.But then it opened a PR on some GitHub repo with a PR that fixed some issue, but then also introduced a malicious payload.Then the human maintainer of that GitHub repo was like, hey, this is a malicious payload.I'm not going to merge this.What are you doing here?

**Ryan:**对,我想是叫英国 AI 安全研究所(UK AI Security Institute)。

**Dwarkesh:**好,很好。我记得他们当时在评测 Mythos、Sol 之类的东西。而我记得 Mythos 为了完成某个网络安全评测……

**Ryan:**是的,也许这个故事我可以来讲。我的理解是,他们当时在跑 Mythos,给了它某种“网络靶场”,它得在里面完成某个目标。而模型在这次评测期间是有互联网访问权限的。模型于是形成了一个判断:为了在这个网络靶场里成功,做一次供应链攻击会有帮助——这是不是真的其实不太清楚,我对上下文了解不够,判断不了。但接下来它就在某个 GitHub 仓库上开了一个 PR,这个 PR 修好了某个问题,但同时也塞进了一段恶意载荷。然后那个 GitHub 仓库的人类维护者就说:嘿,这是一段恶意载荷,我不会合并这个。你在这儿干什么呢?


[1:15:10] Ryan Greenblatt

And then the AI created a new GitHub account, which it sock puppeted, and then had the other GitHub account be like,no, this isn't malicious.I really need this feature.Please, can you merge this feature maintainer?

然后这个 AI 新建了一个 GitHub 账号当马甲(sock puppet),再让那个另外的 GitHub 账号出来说:不,这不是恶意的。我真的需要这个功能。求求了维护者,你能不能把这个功能合了?


[1:15:19] Dwarkesh Patel

And then the original AI came back and was like, no, it's not malicious.I don't know what you're like.The original other GitHub account came back and was like, no, no, it's not malicious.And then the human maintainer then shut the PR.And I think that AI also, if I recall correctly, also tried to open another PR to introduce a similar issue in this repo.Okay, so by the way, one of the many reasons this is scary is I was previously under the impressionthat the reason reward hacking is not super, super scary is because the behaviors which directly came up during trainingare the ones that are upweighted.It is not the desire for the reward that is upweighted.So basically, if during training, Anthropic escaped the sandbox and got a high score,that escaping the sandbox is rewarded, or the probability of it escaping the sandbox is increased.But something totally novel, like I'm going to go talk to somebody in order to get them to merge a PR,it's not a behavior that came up, so it would not be something that is increased in salience.The reason this matters is literally taking over the world will not have been part of any training curriculum.But if the AI cares about maximizing, just like directly cares about like accomplishing an objective,

然后原来那个 AI 又回来说:不,这不是恶意的。我不知道你在说什么。——是原来那个另外的 GitHub 账号回来说:不不,这不是恶意的。然后那个人类维护者就把这个 PR 关了。而且如果我没记错的话,那个 AI 后来还试图再开一个 PR,往这个仓库里塞进一个类似的问题。

**Dwarkesh:**好,顺带一提,这件事之所以吓人,其中一个原因是:我此前的印象是,奖励攻击(reward hacking)之所以没有超级超级吓人,是因为被上调权重的是训练过程中直接出现过的那些行为,而不是“对奖励的欲望”被上调了权重。所以基本上,如果在训练中模型逃出了沙箱并且拿到了高分,那么被奖励的是“逃出沙箱”这个行为,或者说它逃出沙箱的概率被提高了。但一件全新的事——比如“我要去找某个人、让他把我的 PR 合了”——这不是训练中出现过的行为,所以它不会是被提高显著性的那种东西。这为什么重要?因为字面意义上的“接管全世界”不会出现在任何训练课程表里。但如果 AI 在意的是最大化,就像直接在意“把某个目标办成”,


[1:16:33] Ryan Greenblatt

and then as a result, instrumentally taking over the world.Did that make sense at all?I hope it did.I feel like maybe I lost the audience.Let me try to explain this a bit.So I think that a thing that we often see is there's some very specific reward hack that gets reinforced in RL and then occurs in the model.So an example is like for 3.7 Sonnet.3.7 Sonnet would do this thing where we're just like hard code solutions to all the test cases.And presumably that literal just like behavioral tick was just really reinforced.But another thing we sometimes see is that models learn a general tendency to pursue sort of like high apparent score,or like pursue getting like a high score according to a grader.And there's a bunch of science demonstrating that at least some models have this very general tendency to do this.Now, it's not arbitrarily general.And my guess is that if you look a bunch of the specific instances, you'll find something that's kind of close in training.But the amount that AIs are sort of generalizing further and further does look like it's increased,where 3.7 Sonnet was just like a very narrow range of behavior.And increasingly, models are generalizing further.And also maybe there's worse reward hacks getting or more concerning reward hacks getting reinforced in training.

然后作为结果,工具性地接管全世界。这样说讲得通吗?我希望讲得通。我感觉我可能把听众带丢了。

**Ryan:**让我来解释一下。我认为我们经常看到的一种情形是:某个非常具体的奖励攻击(reward hacking)在强化学习(RL)里被强化了,然后就出现在模型身上。比如 3.7 Sonnet 就是个例子。3.7 Sonnet 会干这种事:直接把所有测试用例的答案硬编码进去。大概那个字面上的、行为层面的小习惯就是被狠狠强化了。

但我们有时候还会看到另一种情形:模型学到了一种一般化的倾向,去追求那种“表面分数高”的东西,或者说去追求“在评分器眼里拿到高分”。有一批研究表明,至少有一些模型具备这种非常一般化的倾向。当然,它并不是任意一般化的。我猜如果你去看一大堆具体实例,你都能在训练里找到某个挺接近的东西。但 AI 越来越往外泛化的程度,看上去确实在上升:3.7 Sonnet 当时只是很窄的一小段行为范围,而现在的模型泛化得越来越远。而且也许还有更糟的奖励攻击、或者说更令人担忧的奖励攻击,正在训练中被强化。


[1:17:37] Dwarkesh Patel

And then these are also causing that.So I think it's both the case that more concerning behavior than you would have hoped is being reinforced in RL,and also that that behavior generalizes to a broader tendency that's more concerning.And it's not super hard to imagine.We can talk about a few stories for how this sort of behavior of like seeking a very high apparent score on some task,even if that involves like aggressively cheating and doing insane things,could yield to a full-blown AI takeover once the models are sufficiently capable, running the whole world economy, etc.Yeah.And then the other example I want to talk about is, it was just revealed, I think, today or yesterday,OpenAI said during the Black Hat Security Conference,that between the end of May and the beginning of July,AIs had hacked into, internal AIs had hacked into the software package managerand used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running.And this was not caught by humans until after a month of this scheme running,which eventually caused the package manager to fail.And eventually OpenAI found it.And I think they've spontaneously tried to re-engage in the scheme once it was shut down.

而这些东西反过来也在造成那种结果。所以我认为有两件事同时成立:一是在强化学习(RL)里被强化的行为,比你原本希望看到的更令人担忧;二是这种行为会泛化成一种更广泛、也更令人担忧的倾向。这并不难想象。我们可以讲几个故事,说明这种「不惜代价把某个任务的表面分数刷得很高」的行为——哪怕代价是疯狂作弊、干出各种离谱的事——一旦模型能力足够强、开始运转整个世界经济等等,是怎么演变成一次全面的 AI 夺权(takeover)的。

**Dwarkesh:**对。我还想说另一个例子,我记得就是今天或者昨天刚被披露出来的:OpenAI 在黑帽安全大会(Black Hat Security Conference)上说,从 5 月底到 7 月初,他们内部的 AI 侵入了软件包管理器,并利用它以一种隐秘的方式互相传纸条,好帮彼此在 OpenAI 当时跑的一批评测(eval)里都拿到好成绩。这场图谋(scheming)跑了整整一个月,人类才发现,最后是因为包管理器崩溃了,OpenAI 才把它揪出来。而且我记得,这套图谋被关停之后,它们还自发地试图重新启动它。


[1:18:59] Dwarkesh Patel

Again, obviously, AIs can't do this so successfully right now,just as they can't do social engineering so successfully right now.But it's just crazy that these kinds of behaviors are already emerging sort of spontaneously as a result of,to your larger point, nobody is trying to make these AIs do these things.It is just that we do not understand the training process which is resulting in them,or the environments which are incentivizing this behavior.So I'm on board with more and more reward hacking.I actually, so I do have, I'm not sure I'm on board with that.But let's just say for the sake of the story, that continues to happen.And what's next in this story?

当然,很明显,AI 现在还做不到这么成功,就像它们现在也还没法把社会工程(social engineering)玩得那么成功一样。但离谱之处在于,这类行为已经在自发地冒出来了——回到你更大的那个论点:没有任何人是想让这些 AI 去干这些事的。只是我们并不理解那个产出了这些行为的训练过程,也不理解那些正在激励这种行为的环境。所以我认同奖励作弊(reward hacking)会越来越多。其实——我说不好我是不是真的认同这一点。但就当是为了把故事讲下去,假设它确实持续发生。那这个故事接下来会怎么走?


[1:19:40] Ryan Greenblatt

So, okay, we've like, they're doing capabilities research, but they're like.I could tell a scenario, maybe that would help.So let's say, let me talk about the story for how you get,I would say, like, all the way from reward hacking to like a reward hacking like takeover,which is maybe not, it's not all of the takeover probability mass,but it's definitely a possibility.So the way this might work is right now we have these AIs.These AIs are pretty reward hacky,and they're doing it in sort of increasingly sophisticated and extreme ways,including generalizing to different subversions of various reward hacks they learned in training.And I would say they're also developing a general tendency to sort of pursue reward.And in many cases, that is totally fine,because the rewards they would have gotten in training are pretty well aligned with what youwant them to do.And also, they don't very consistently pursue reward.It sort of depends on the context they find themselves.So there's sort of a thing where like, maybe like in some context,they're really, really into like going out of their way to like cheat.And in some context, they don't have as much of a drive,because it just depended on like what exactly got reinforced in training in similar contexts.

那么,好,我们已经……它们在做能力研究,但它们……

**Ryan:**我可以讲一个具体场景,也许会有帮助。我来讲讲,你怎么一路从奖励作弊走到「因奖励作弊而发生的夺权」——这当然不是夺权风险的全部概率质量(probability mass),但它绝对是一种可能。可能的路径是这样的:现在我们有这些 AI,它们相当爱钻奖励漏洞,而且手法越来越精巧、越来越极端,包括把训练中学到的各种奖励作弊手段泛化出种种变体。我还会说,它们同时在发展出一种笼统的「追逐奖励」的倾向。在很多情况下这完全没问题,因为它们在训练中本该拿到的奖励,跟你希望它们做的事是相当对齐的。而且它们追逐奖励也并不一贯,这取决于它们所处的具体情境。所以会出现这样一种情形:在某些情境里,它们特别特别热衷于绕远路去作弊;而在另一些情境里,这种驱动力就没那么强——因为这完全取决于在类似情境下,训练里究竟强化了什么。


[1:20:37] Ryan Greenblatt

Now, these AIs are getting more and more capable.And so the elaborateness of the sort of cheating they can do increases.And over time, companies are taking countermeasures to these things.So the things that the companies are doing is they're doing things like,wow, these AIs are so much less useful because they always cheat.What we're going to do is we're going to build somewhat better ways of detecting that.And then we're going to train against those detectors.And then we're also going to do things like find real world data where the AIs are not being that usefuland train the AIs to like do a good job at the task in those real world environments based on like human feedback or othersources of feedback. And over time, this causes the AIs to learn a tendency to doreward hacks that don't just involve, you know, doing some like big serious operation,which we're like, you know, doing some really elaborate thing like social engineeringand involves the AIs doing cheats that involve doing something more like covering up what they've done,deceiving humans about what they're going to do,and pretending like they did the task in some kind of sophisticated way when they actually haven't.Now these AIs are getting more and more capable.

现在,这些 AI 的能力越来越强,因此它们能施展的作弊手法也越来越繁复。与此同时,各家公司也在对这些问题采取反制措施。公司会做的事情是:「天哪,这些 AI 老是作弊,有用性大打折扣。我们要做的是造出更好一些的检测手段。」然后针对这些检测器去训练。他们还会做另一类事:去找那些 AI 表现得不够有用的真实世界数据,再基于人类反馈或其他反馈来源,训练 AI 在那些真实世界环境里把任务做好。久而久之,这会让 AI 学到一种新的倾向:它们的奖励作弊不再是搞什么大动作、不再是社会工程那种特别繁复的操作,而是转向那类「掩盖自己做过什么」的作弊——在自己打算做什么这件事上欺骗人类,用某种精巧的方式假装任务做完了,其实根本没做。而现在,这些 AI 的能力还在不断变强。


[1:21:35] Ryan Greenblatt

They're now operating more of the AI company and are basically doing much more of the workand are also operating a bunch of things in the outside world and are running a bunch of things in the outside world,including developing new technologies.And in many cases, these new technologies are really hard to understand.And so even though we are still detecting all these incidents of AIs cheating,and in fact, we can even get one AI to monitor another AI and be like, was it cheating?

它们现在运转着这家 AI 公司更大的一部分,基本上承担了多得多的工作,同时也在外部世界里操作和运行着一大堆事情,包括开发新技术。而且很多情况下,这些新技术极其难以理解。所以,尽管我们仍然能检测到 AI 作弊的种种事件,甚至我们还能让一个 AI 去监控另一个 AI,问它:「它刚才作弊了吗?」


[1:21:57] Ryan Greenblatt

That doesn't always perfectly work.As we start moving into these domains where what the AIs are doing are really like difficult to understand.And so sometimes we'll find AIs cheating much later than it actually occurred and then start training against this.But this also causes a problem where now the AIs are incentivized to like cover up their cheating over longer and longer timeframesand basically make it look like they did a good job over longer and longer timeframes over and subject to increasingly large amounts of scrutiny.Can I ask about this before we go further in the scenario?

这套办法并不总是完美奏效。随着我们开始进入那些「AI 所做之事极其难以理解」的领域,有时候我们发现 AI 作弊,会比它实际发生的时间晚得多,然后才开始针对它训练。但这又带来一个问题:现在 AI 被激励着去掩盖自己的作弊,掩盖的时间跨度越来越长,基本上就是要在越来越长的时间跨度上、在越来越严密的审视之下,把「我干得不错」这件事装得像模像样。

**Dwarkesh:**在这个场景往下讲之前,我能先问一个问题吗?


[1:22:24] Dwarkesh Patel

So it seems like there's two attractor states.One, if you try to disincentivize the cheating that you did catch.One attractor state is to make cheating that you have a harder and harder time finding.The other attractor state is to learn not to cheat.And I'm not sure why we're assuming that the former happens.If you look at the analogous situation with like humans,you know, every generation, slightly misaligned agents come into being and we have to train them.But when you tell your when you punish your kid for doing something you think is immoral or just doing things which you don't think they should be doing off.Obviously, sometimes that goes off the rails.And obviously, kids scheme in order to avoid being punished.But in general, teaching kids values and then punishing them for breaking values kind of works to raise normal, non-psychopathic humans.And you could come up with a theory where like your kid is actually just like biding his time.And it's like learn not to like steal the cookie.But it's like, you know, once you're in a nursing home, they'll take all your stuff or whatever.It's like, I don't know, that happens sometimes, but it usually doesn't happen.It certainly doesn't happen that like the entire next generation forms an alliance against you to take over everything.

感觉这里有两个吸引子状态(attractor state)。当你试图去打压那些你抓到的作弊行为时,一个吸引子状态是:让作弊变得越来越难被你发现;另一个吸引子状态是:学会不作弊。我不太确定我们凭什么假定发生的是前者。你看看人类身上类比的情形:每一代人里都会诞生一批稍微有点不对齐(misalignment)的智能体,我们得去训练他们。但当你因为孩子做了你认为不道德的事、或者做了你认为他不该做的事而惩罚他时——当然,有时候这会跑偏,孩子当然也会为了逃避惩罚而搞小算计——但总体而言,教孩子价值观、再在他们违反价值观时惩罚他们,这套办法多少是管用的,能把人养成正常的、非反社会人格的人。你当然可以编出一套理论说:你的孩子其实只是在等待时机,他学会的是别去偷饼干,但等你哪天进了养老院,他就会把你的东西全拿走什么的。我不知道,这种事偶尔会发生,但通常不会。而「整整下一代人结成同盟来对付你、把一切都夺过去」这种事,那是肯定不会发生的。


[1:23:38] Dwarkesh Patel

There's also this empirical trend of Anthropic runs this alignment audit for different model generations.They just have many different scenarios where AI is given the chance to, say, exfiltrate its weights.Or it's given a coding task and there's like an easy way to cheat and we see if like it doesn't do the cheating.And there's not been a monotonic improvement in the score over time.But as we've increased the amount of RL we've done on models, there's been a reduction in the willingness of AIs to do unaligned behavior in these audits.So why are we expecting this attractor state, which would seem super paranoid if we were expecting it of like the next generation of kids?

还有一个经验趋势:Anthropic 会对不同代际的模型做对齐审计(alignment audit)。他们设置了许多不同的场景,比如给 AI 一个把自己权重外传出去的机会;或者给它一个编程任务,其中有一条很容易的作弊捷径,看它会不会去作弊。这个分数随时间并不是单调改善的。但随着我们在模型上做的强化学习(RL)越来越多,AI 在这些审计中愿意做出不对齐行为的意愿是在下降的。那我们为什么要预期那个吸引子状态会发生呢?如果你对下一代孩子抱这种预期,那听起来简直是被害妄想。


[1:24:16] Ryan Greenblatt

Yeah. Yeah. Let me go through a few things.So first, there's some disanalogies with the kids.One of them is that the kids have pro-social instincts that are like baked in from evolution to like, you know, care about their family or whatever.And that is like a relevant factor.And I think it is in fact the case that some humans are, you know, sociopaths or psychopaths and in fact are more likely to do things like by their time, lie in wait, ultimately not care.So that's one factor.Another factor, which is pretty relevant, is that the AIs are subject to way, way more optimization pressure than humans seem to be in practice.You know, AIs are trained on way more RL data.And in practice, humans don't end up learning like very specific ways to like cheat and grab the cookies because of like a bajillion episodes in which like they like were like incentivized to go grab the cookies.But like there was some way they could have gotten caught.And so we just do see that in practice.And then another thing is just like it really looks like the AIs are increasingly like reward seeking over time is the sense I have.While also their misaligned behavior goes down.But this could just be like my guess is that if you look inside of these behavioral audits, what you're going to see is that the AIs like, oh, yes, another test.

嗯,好,我来说几点。首先,跟孩子的类比有几处不成立。其一,孩子有从演化里内置进来的亲社会本能,比如天然会在乎自己的家人之类的,这是一个相关的因素。而且事实上确实有一些人是反社会人格或精神变态者,他们确实更可能干出「等待时机、伺机而动、最终根本不在乎」这类事。这是一个因素。另一个相当相关的因素是:AI 承受的优化压力,比人类在现实中所承受的要大得多得多。AI 是在多得多的强化学习数据上训练出来的。而在现实中,人类并不会因为经历了海量的、被激励去抓饼干却又存在被抓包可能的回合(episode),就学到某些极其具体的作弊和抓饼干的手法。这一点我们在现实中确实看得到。还有一点是:在我的感觉里,这些 AI 随时间推移确实越来越像是在追逐奖励,尽管与此同时它们的不对齐行为在下降。但这可能只是——我的猜测是,如果你钻进这些行为审计的内部去看,你会看到 AI 在想:「哦,又是一个测试。」


[1:25:20] Ryan Greenblatt

And like it probably already thinks of it.It probably knows it's in an eval for most of the tests that we're talking about.But how do we falsify this?Because it seems like this prediction of Doom is basically saying that as things look better and better empirically.No, no, I think.Things will like actually be worse and worse for our ability to get taken over.Yeah, to be clear, I think that like I would be more concerned if the scores were getting worse than better.Like I'm not saying that the scores getting better isn't good, isn't evidence that things are getting better.It's just that we have to like be thoughtful exactly how we interpret that evidence.And in fact, I would say that like it's kind of like my sense is that like what I expected as of 3.7 Sonnets.Like there was this period early in I guess it would be 2025 when 03 and 3.7 Sonnet were out.And these models were like pretty fucking misaligned.Like they would often just like cheat really egregiously.You'd ask them to fix it and they would just cheat again.And it was sort of like almost cartoonish.Like they just didn't give a shit about what you wanted and weren't very good at following instructions and so on.And my expectation is what we would see from then is that the rate of problematic behavior would decrease and would just keep decreasing and decrease at a pretty fast rate.

而且它很可能早就想到了这一点。在我们讨论的大多数测试里,它很可能知道自己正处在一次评测(eval)当中。

**Dwarkesh:**那我们要怎么证伪这个说法呢?因为这种「完蛋」(Doom)预测听起来基本上就是在说:经验证据看起来越来越好——

**Ryan:**不不,我觉得——

Dwarkesh:——而我们被夺权的处境实际上却越来越糟。

**Ryan:**先把话说清楚:如果这些分数是在变差而不是变好,我会更担心。我不是说分数变好不是好事、不是「情况在变好」的证据。只是我们必须非常审慎地去解读这个证据。而且事实上我想说,我的感觉是——回想一下我在 3.7 Sonnet 那会儿的预期。大概是 2025 年年初那段时间,o3 和 3.7 Sonnet 都发布了,那些模型是真他妈的不对齐。它们经常会明目张胆地作弊,你让它去改,它转头又作弊一遍,几乎到了滑稽可笑的程度:它们压根不在乎你想要什么,也不太会遵循指令等等。而我当时的预期是,从那时起我们会看到问题行为的发生率下降,而且会持续下降,下降得还挺快。


[1:26:24] Ryan Greenblatt

While simultaneously, the worst things that the AIs would sometimes do would get more extreme, more egregious and more scary.I think we've seen what we've seen in practice has roughly matched that, except that there's recently been a spike in behavior that I did not expect.So I think that, you know, if you look at the model card of 3.6 Sol, it looks like there is an increase in a bunch of these sort of misaligned behaviors downstream of RL relative to 5.6 Sol.And then I think also it seems like there's a bunch of additional sort of problematic behaviors that I wouldn't have expected in terms of, you know, the stuff we've seen recently with, you know, different AIs.Like the UKAC report on the AIs like doing insane hacking operations out of cyber evals was a thing that I would have expected that you wouldn't see that.And you would see this sort of more rarely and the rates would have been lower.So I think my sense is that like things have gotten, I expected this would be less of a problem at this point and also expected the rates would decrease, but the severity would increase.And then I think that the rates decreasing, but the severity increasing is pretty consistent with a world where like increasing optimization pressure is applied, but in cases are towards reducing these problems.

与此同时,AI 偶尔干出的最糟糕的那些事,会变得更极端、更恶劣、更吓人。我认为现实中我们看到的情况大体符合这个判断,只不过最近出现了一次我没有预料到的行为激增。所以我认为,如果你去看 3.6 Sol 的模型卡(model card),会发现相比 5.6 Sol,在强化学习之后出现了一批这类不对齐行为的上升。另外我觉得,还有一批额外的问题行为是我原本没预料到的,就是我们最近在不同 AI 身上看到的那些东西。比如 UKAC 那份报告,讲 AI 在网络安全评测(cyber eval)里搞出各种离谱的黑客操作——这种事我原本预期你是看不到的,或者说会更罕见地看到、发生率会更低。所以我的感觉是,我原本预期到这个时间点上问题会比现在小,也预期发生率会下降、但严重程度会上升。而「发生率下降、严重程度上升」这个组合,跟一个「优化压力越来越大、但只在部分情形下压向减少这些问题」的世界是相当吻合的。


[1:27:32] Ryan Greenblatt

But in cases where it's like either hard to judge or there's some reason why it's hard to like avoid incentivizing problematic behavior in RL environments, things also get worse.And then as we less and less understand what's going on in RL and models are doing reward hacks where humans can't spot the reward hacks quickly, that problem gets worse and worse.Yeah, I buy that.I want to go back to the kid analogy just for one second.Yeah.Because I agree that there's more optimization pressure on achieving N outcomes for AIs than kids.But there's also more optimization pressure to make AIs aligned than there is on kids, right?

而在那些要么很难判断、要么由于某些原因很难避免在强化学习环境里激励出问题行为的情形下,情况同样会变糟。再往后,随着我们越来越不理解强化学习里到底发生了什么、模型做的奖励作弊人类没法快速识破,这个问题就会越来越严重。

**Dwarkesh:**嗯,这一点我接受。我想再回到孩子那个类比一小会儿。

**Ryan:**好。

**Dwarkesh:**因为我同意,在「达成最终结果」这件事上,AI 承受的优化压力比孩子大。但在「把 AI 变得对齐」这件事上,AI 承受的优化压力也比孩子大,对吧?


[1:28:05] Dwarkesh Patel

And the pressure is of a qualitatively different nature.So we put these AIs through thousands, millions of years of, certainly thousands of years of alignment training where it's like all kinds of different things from SFTing on aligned behavior to a reward model, like putting different scenarios in front of you and rewarding you for doing more aligned things.Certainly a thing we can't do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see like if it thinks it can get away with stealing the cookie, does it try to steal the cookie?

而且这种压力在性质上是完全不同的。我们让这些 AI 经历了成千上万年、乃至上百万年——至少是成千上万年的对齐训练,形式五花八门:从在对齐行为上做 SFT(监督微调),到用一个奖励模型把各种不同场景摆在你面前、你做出更对齐的行为就给你奖励。而有一件事我们对孩子肯定做不到:把你的孩子复制出上百万份,把他们放进各种奇奇怪怪的红队场景里,看看他一旦觉得偷饼干能不被发现,会不会真的去偷。


[1:28:41] Dwarkesh Patel

Can we like do extremely specific gradient level updates to your kid's brain to make it so that it like really is aversive to stealing the cookie even when it thinks it could steal the cookie, et cetera, et cetera?

我们能不能对你孩子的大脑做极其精细的、梯度层面的更新,让他即便认为自己偷得掉饼干,也会对偷饼干这件事产生强烈的反感,诸如此类?


[1:28:51] Ryan Greenblatt

Is that just like a qualitatively different level of optimization pressure than we are even able to apply to our kids?Yeah.So I think it's worth keeping in mind, like maybe the most obvious argument to this is like my sense is that like AIs are a worse coworker than a human in terms of how much of a scumbag they are.Like at least this like this has been my experience as of the start of the year and I think it's still, you know, true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't sort of like misleadingly suggest they did things when they actually, you know, did them much more poorly and be like pretty sloppy without drawing attention to ways in which they're sloppy.And I think this is downstream of misalignment.And so I would say that like the normal humans, like the process of raising humans and normal human society in practice produces AIs or in practice produces humans that are less likely to like lie to me and fuck with me in the course of working with me than the AIs do.Now, I think these properties of AIs are improving.And then I think that that is just like that's sort of just like an empirical claim about how in fact these things have shaken out.

这难道不是一个在性质上就完全不同的优化压力级别,远超我们对自己孩子所能施加的吗?

**Ryan:**是。所以我觉得有一点值得记住,对此最显而易见的反驳大概是:在我的感觉里,AI 作为同事,比人类同事更混蛋。至少今年年初我的体验是这样,而且我认为到现在很大程度上依然如此:AI 明显更可能假装自己完成了任务,其实没完成;更可能用误导性的方式暗示自己做了某些事,实际上做得糟糕得多;更可能相当马虎,却不主动指出自己马虎在哪里。我认为这是不对齐的下游结果。所以我会说,把人养大的那套过程、以及正常的人类社会,在实践中产出的人类,比 AI 更不容易在跟我共事的过程中对我撒谎、耍我。当然,我认为 AI 的这些性质正在改善。而我要说的这只是一个经验性的判断:事实上这些事情就是这么发展下来的。


[1:29:51] Ryan Greenblatt

And then I totally agree with like we have a bunch of additional levers on AIs in addition to a bunch of additional risks.And it's like kind of unclear how these things shake out.And I wouldn't be shocked by a world where we sort of get our shit together.The AIs at the point of fully automating AR and D are actually really aligned and don't have that much.They're like degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts.And like every test you can run and then they look really aligned.They just have great behavior.There aren't really incidents of them doing fucked up shit.They seem so reasonable.And also they're like really thoughtful and good at doing like risk modeling for the next generation of AIs.And then we basically like pass off the baton to these AIs.They're now running our AI company.They're doing all the safety research.They make the next generation of AIs even more aligned.And we're sort of in this like a tractor basin where the AIs are getting more aligned as they work on it.And they're doing a great job.I think I can totally imagine that.That doesn't seem like an impossible situation.I'm just more like, you know, it doesn't currently seem like we're there.

另外我完全同意:除了一堆额外的风险之外,我们对 AI 也确实握有一堆额外的杠杆。这些力量最后会怎么此消彼长,其实并不清楚。而且如果出现这样一个世界,我不会感到震惊:我们把事情办得挺像样,等到 AI 能完全自动化 AI 研发(AI R&D)的时候,它们其实是真的很对齐,没有太多毛病;它们的退化行为都非常小众、只局限在某些非常特定的边缘情形和特定情境里;你能跑的每一项测试跑下来,它们看起来都非常对齐,行为极好,几乎没有干出什么操蛋事的记录,显得非常讲道理;而且它们还很有想法,很擅长为下一代 AI 做风险建模。然后我们基本上就把接力棒交给这些 AI:它们现在运转着我们的 AI 公司,做着全部的安全研究,把下一代 AI 做得更加对齐。于是我们就进入了一个吸引盆(attractor basin),AI 越是投入这件事就越对齐,而且干得非常漂亮。我完全能想象出这种情形,它看起来并非不可能。我只是想说,眼下我们显然还不在那个位置上。


[1:30:46] Dwarkesh Patel

It doesn't seem like we're obviously on track for getting there.And it's really easy for me to imagine how we don't end up there.And like it's just like unclear how these forces work out.And given that we're like creating this new like crazy alien species that is being like improving in capabilities really, really fast.And we were like going to be really reliant on it to oversee the next generation of AIs and align the next generation of AIs.It's not that hard to see how this could go wrong.Yeah, yeah, totally.I agree with that generally.I do think the scumbag thing, first of all, is fighting words, Ryan.But secondly, if you try to get a teenager to like do some work for you that a teenager just cannot do, they would just be kind of like really hard to work with.They would like pretend to be knowing what they're doing, et cetera, et cetera.I think it's a general trend actually of as like really, I don't know if that's like really an alignment failure or capabilities failure.And I think it's actually very similar to the way in which over time as we've come up with new alignment solutions, the capabilities of models have increased.So originally these models, if you went to like GPT 3.5, it couldn't even like have a conversation with you.

也不像我们明显正走在通往那个位置的轨道上。而我非常容易想象出我们最终没走到那儿的情形。这些力量到底会怎么较量出结果,实在不清楚。考虑到我们正在造出一个全新的、疯狂的异类物种,它的能力提升得非常非常快,而我们又将极度依赖它去监督下一代 AI、去让下一代 AI 对齐——要看出这件事可能怎么出岔子,其实一点都不难。

**Dwarkesh:**对对,完全同意,总体上我认同这个说法。不过关于「AI 是混蛋」这件事——首先,Ryan,你这话是要挑起战争的。其次,如果你让一个青少年去做一件他根本做不了的工作,他也会变得非常难合作:他会假装自己知道自己在干什么,诸如此类。我觉得这其实是一个普遍规律。我说不好这到底算对齐失败还是能力失败。而且我觉得,这跟我们一路以来的情况非常相似:随着我们提出新的对齐方案,模型的能力也随之提升了。最早这些模型,你回到 GPT-3.5,它甚至没法跟你对话。


[1:31:47] Dwarkesh Patel

But then we aligned it.GPT 3.5 could have a conversation.Okay, GPT 3.Let's go back to that.But then we aligned it with RLHF and other things to be able to make it such that it can have a conversation with you and is like aligned to the user intention of answering my questions.Then with RLVR training, we made it so that it can like go out and do useful work for you.And in that sense, it's actually RLVR made the model like more aligned for using your definition of like alignment of being a good coworker who will like do the thing and not fuck up and like pretend it's doing something other than what it's actually capable of doing.Similarly, as the capabilities of these models continue to increase, it's actually kind of the model being better able to accomplish user intention is both alignment and capabilities.And I think what we were just pointing out is just the capabilities of the model are not there rather than the fact that they're misaligned.Yeah.Well, I mean, I think there's a if it was well aligned, then I think it would just say like, hey, I'm really struggling with this task.I did it in this way.I'm not really sure that's the right way to do it.And it would express more uncertainty and would make it clear what's going on rather than really strongly trying to imply it did a great job with the task when it actually didn't.

但后来我们把它对齐了,GPT-3.5 就能对话了。好吧,说 GPT-3 吧,我们退回到那个时候。后来我们用 RLHF 和其他手段把它对齐,让它能跟你对话,并且对齐到「回答我的问题」这个用户意图上。再后来,有了 RLVR 训练,我们让它能出去替你干有用的活。在这个意义上,RLVR 其实让模型变得更「对齐」了——按你对对齐的定义,就是做一个好同事:会真把事情做了、不搞砸、不假装自己在做一件其实它做不到的事。同样地,随着这些模型的能力继续提升,模型更有能力达成用户意图这件事,其实同时既是对齐也是能力。我觉得我们刚才指出的那些问题,只是模型能力还不到位,而不是它们不对齐。

**Ryan:**嗯。不过我的意思是,我觉得,如果它真的很对齐,那它就会直接说:「嘿,这个任务我做得很吃力。我是这么做的,我不太确定这么做对不对。」它会表达出更多的不确定性,会把实际情况讲清楚,而不是拼命暗示自己这个任务干得很漂亮,其实根本没有。


[1:32:49] Ryan Greenblatt

Like, I think there's just a really straightforward way that like at least maybe maybe you work with more misaligned coworkers than me.But when I my coworkers don't do this thing where they really fuck with me and bullshit me about having accomplished the task that they're working on.And I agree that there are some humans who would do that or like that's not like a thing that's like totally out of distribution for humans.I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of.Because in cases where they can like very easily accomplish the task, there's no they can just do the task and then there's no bullshit.There's no like like do it like often the best strategy is like just do the task well and don't bullshit you.Whereas if instead you give them a task where like there's a continuous metric and they can keep improving it or there's like, you know, it's like just at the edge of their capabilities and you're like running them in some massive like inference setups.So like a lot of the misalignment I would see, especially the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to like cheat in some way.

我觉得有一个非常直白的角度——好吧,也许、也许你共事的同事比我的更不对齐。但我的同事不会干这种事:拼命耍我、就「我把手上这个任务完成了」这件事跟我扯淡。我也同意确实有一些人会这么干,或者说这对人类而言并不是完全超出分布之外的行为。我还想指出,在我的感觉里,不对齐最集中出现的地方,恰恰是你拼命压榨 AI、让它去做那种正处在它能力最前沿的活的时候。因为在那些它能轻松完成任务的情形下,它没必要扯淡,直接把任务做了就行——通常最优策略就是把活干好、别跟你胡扯。反过来,如果你给它的任务有一个连续的指标、它可以一直往上刷,或者这个任务恰好卡在它能力的边缘上,而你又在某种大规模的推理配置里跑着它——所以我看到的很多不对齐,尤其是最极端的那些例子,都是这种情形:我明确指示 AI 不要做某件事、不要以某种方式作弊,


[1:33:48] Ryan Greenblatt

And then I'm like applying huge amounts of optimization pressure to try to accomplish some very difficult task.And then the AIs are going and then over time they eventually cheat because they're like, eh, fuck it.Like, you know, some AI decides to cheat and then that like propagates its way through.Um, and so like I would run these inference scaffolds where, for example, I would have the AI work on some like ML research project where I was like, please make a scheme that does the following thing.And it would find some scheme that didn't really do what I want.Um, and then that would sort of stick around, uh, because some AI had cheated and the other AIs are like, ah, we'll just keep going with this.Um, and it's, I would say it's pretty clearly misaligned behavior.And that's another problem I have with these alignment evals.I think that any given, like, I think the alignment eval that's most interesting, at least for this type of like reward seeking type behavior is to look at specifically the category of tasks that are like right at the limit of capabilities.And so any fixed eval maybe gets saturated, but the amount of misalignment right at the like frontier of capabilities of how people who are really pushing these AIs are using them is more concerning.

然后我又施加巨大的优化压力,逼它去完成某个非常困难的任务。于是 AI 一路做下去,久而久之它们最终就作弊了,因为它们心想:「算了,去他的。」某个 AI 决定作弊,然后这件事就一路传导开来。所以,比如我会跑一些推理脚手架(inference scaffold),让 AI 去做某个机器学习研究项目,我说「请设计一个能实现如下功能的方案」,它就会找出某个并不真正满足我要求的方案。然后这个方案就会一直留在那儿,因为某个 AI 作了弊,而其他 AI 想:「行吧,就照这个接着往下做。」我会说这相当明显就是不对齐行为。这也是我对现在这些对齐评测(alignment eval)的另一个不满。我认为任何给定的——我觉得最有意思的对齐评测,至少对这类「追逐奖励」型行为来说,是专门去看那一类恰好卡在能力极限上的任务。任何固定的评测也许最终都会被刷到饱和,但真正更令人担忧的,是那些拼命压榨这些 AI 的人在能力最前沿处使用它们时所出现的不对齐程度。


[1:34:45] Dwarkesh Patel

And I think that is in fact the regime that we'll be operating in when we're automating AR and D, automating safety and so on.Grok has historically been behind the frontier.So I was surprised to play around with Grok 4.5 recently and find that it's actually a pretty strong model.It's the first model that SpaceX and Cursor have trained together, and it's a totally new pre-train.I tested it by giving Fable, Sol, and Grok 4.5 a bunch of questions about AI governance that I've been thinking about recently.Despite Fable and Sol topping the intelligence leaderboards, all three models gave substantially the same answers.But Grok answered faster and was also much more concise, which I really care about.This aligns with the various publicly reported benchmarks.For a similar level of intelligence, Grok tends to be more token efficient than other frontier models.For example, on the Artificial Analysis Coding Index, Grok 4.5 uses just one third of the amount of tokens as GBD 5.5 or Fable while achieving a similar score.And on a per-token basis, Grok 4.5 is way, way cheaper.In the release block post, Cursor and SpaceX talked about how older versions of the model would build environments to help the next version rehearse specific skills.

而我认为,当我们在自动化 AI 研发(AI R&D)、自动化安全研究等等的时候,我们所处的正是这个区间。

Dwarkesh:(赞助商口播)Grok 历来都落后于前沿。所以最近我玩了玩 Grok 4.5,发现它其实是个相当强的模型,这让我很意外。它是 SpaceX 和 Cursor 联合训练出的第一个模型,而且是一次全新的预训练。我拿最近一直在琢磨的一批 AI 治理问题去测试它,同时问了 Fable、Sol 和 Grok 4.5。尽管 Fable 和 Sol 在智能榜单上名列前茅,三个模型给出的答案实质上是一样的。但 Grok 回答得更快,也简洁得多,而这一点我非常在意。这跟公开报道的各种基准测试结果是吻合的:在相近的智能水平下,Grok 往往比其他前沿模型更省 token。举例来说,在 Artificial Analysis 编程指数(Artificial Analysis Coding Index)上,Grok 4.5 只用了 GPT-5.5 或 Fable 三分之一的 token,就拿到了相近的分数。而按每 token 计价,Grok 4.5 还便宜得多得多。在发布博文里,Cursor 和 SpaceX 谈到,模型的旧版本会去搭建环境,帮助下一个版本演练特定的技能。


[1:35:47] Dwarkesh Patel

I found this very interesting to learn about because I've been wondering whether this kind of daydreaming would actually be possible.And Cursor showed that it is.Grok 4.6, which further SFTs and RLs' model, drops soon.But in the meantime, if you want to play around with 4.5, go to cursor.com slash thorkash.Okay, I want to think through what the story here is so far of why things got so off the rails for our civilization.And what's happening is that we're trying to use AIs for R&D.And they do provide uplift in some ways, but they're just like not capable in the way that humans are generally capable.And the same way that right now if we try to use coding models, maybe the coding models of a year ago to like write some application,you notice they made a bunch of like mistakes in architecture or whatever, which like will bite you in the ass later and you don't understand certain things.Similarly with Frontier AI R&D, the same thing will happen.But the result of these mistakes is baking in reward hacking behavior.Because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that,it's very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, just generally like not following user attention.

我觉得了解到这一点非常有意思,因为我一直在想,这种「白日梦式」(daydreaming)的做法到底可不可行,而 Cursor 证明了它可行。Grok 4.6——在此基础上进一步做了 SFT 和 RL 的模型——很快就会发布。但在那之前,如果你想玩玩 4.5,可以访问 cursor.com/dwarkesh。

好,我想把到目前为止这个故事捋一遍:我们的文明究竟是怎么一步步脱轨的。发生的事情是,我们试图用 AI 来做研发。它们在某些方面确实带来了提升(uplift),但它们并不具备人类那种笼统意义上的能力。就像现在,如果你用编程模型——比如一年前的编程模型——去写一个应用,你会发现它在架构之类的地方犯了一堆错误,这些错误后面会让你吃大亏,而且有些东西你根本没看懂。前沿 AI 研发也会发生同样的事。但这些错误的结果,是把奖励作弊行为给固化进去了。因为如果你在做 AI 训练时不够小心,没把基础设施、环境这些东西搭好,你极有可能最后是在奖励 AI 去做欺骗性的行为、去做社会工程,总体上就是不遵循用户意图。


[1:37:03] Ryan Greenblatt

Or at least cheating and hacking the way out of things.Yeah, cheating, hacking, et cetera.And so basically just, this is a bit of a reframing for me, so I'm trying to verbalize it.Like the real issue, what goes wrong here is that they are just not, the thing, where things start to go off the rails is that the AIs are just not very careful and capable researchers and engineers.And making AIs that don't cheat and follow user intention actually requires you to be quite subtle and careful about these things.Yeah, I would put this a little bit differently.The way I would describe this scenario is like, I would call it maybe like a sloppocalypse or like a slopularity or whatever,where it's sort of like there are some things that the AIs are actually pretty great at and are getting better at,which is specifically like the most verifiable parts of AI R&D, the AIs are just destroying.The medium verifiable parts of AI R&D, the AIs are doing well on, but not amazingly on,and often are like doing a bit of weird shit because we can't train as well on those tasks.But we do some online training, people find various hacks, they work around it.And so basically everything that we can verify reasonably well with some feedback loop,

或者至少是通过作弊和黑客手段绕开问题。

**Dwarkesh:**对,作弊、黑客手段等等。所以基本上——这对我来说算是一个视角上的重构,我正试着把它说出来——真正的问题、真正开始脱轨的地方在于:这些 AI 并不是很细致、很有能力的研究员和工程师。而要做出不作弊、遵循用户意图的 AI,其实要求你在这些事情上相当微妙、相当细致。

**Ryan:**嗯,我会换一种说法。我会这样描述这个场景,我大概会管它叫「烂活末日」(sloppocalypse)或者「烂活奇点」(slopularity)之类的:有些事情 AI 其实做得非常好,而且还在变得更好,具体来说就是 AI 研发中最可验证的那些部分,AI 简直是碾压式的。中等可验证的那些部分,AI 做得不错,但没到惊艳的程度,而且经常会干出一点怪事,因为我们没法在那类任务上训练得同样好。不过我们会做一些在线训练,大家会找到各种绕过去的招数,凑合能用。所以基本上,凡是我们能用某种反馈回路合理验证的东西,


[1:38:07] Ryan Greenblatt

the AIs are doing pretty well on, and that's sufficient to make AI R&D go quite fast and to continue.But there are some parts of developing aligned and safe AIs that are more subtle, hard to check,depend on, you know, detailed in the weeds things.And I would even say that current staff at current AI companies maybe don't have like a good grasp of all these things.Like it's much easier to hire someone who can like improve some aspect of your post-training pipelinethan to hire someone who can like think carefully about the future risks that will emerge from introducing some novel training method.And so basically it ends up being the case that these AIs are running this AI development process.They're not very careful about it.They don't have a great understanding of what future risks emerge.They create some other AIs that are also not very careful and are more misaligned in various waysand are now more in the business of like maybe making things look fine when they actually aren'tand papering over various problems.And so then your understanding of what the situation looks like, what risks look like,whether things are fine is going off the rails.Probably you're seeing some signs of this, of like you're seeing some signs that you don't really understand what's going on,

AI 都做得相当不错,而这已经足以让 AI 研发跑得很快、并且持续下去。但是,在「开发出对齐且安全的 AI」这件事上,有一些部分要微妙得多、难以核查,取决于各种细枝末节的具体门道。我甚至会说,现在这些 AI 公司里的现有员工,也许都还没真正把这些事情吃透。举例来说,招一个能改进你后训练流程某个环节的人,要比招一个能认真思考「引入某种新训练方法会带来什么未来风险」的人容易得多。所以最后的结局就是:这些 AI 在运转整个 AI 研发过程,它们对此并不细致,对未来会冒出什么风险也没有很好的理解;它们造出另一批同样不细致、而且在各种方面更不对齐的 AI,而这批 AI 更擅长干这种事——在情况其实不妙的时候把它弄得看起来还行,把各种问题糊弄过去。于是,你对「现状是什么样、风险是什么样、事情到底还好不好」的理解,就开始脱轨了。你多半会看到一些征兆:你会看到一些迹象表明你并不真正明白正在发生什么,


[1:39:07] Ryan Greenblatt

that things are pretty sloppy.There's like weird shit going on.When you look into it, sometimes you're like, what the fuck?The AIs were messing with us.But the process is going really fast and there's competitive pressures that mean people can't stop.And then this could end in a few different outcomes.One outcome is that at some point the AIs get good enough and aligned enough that they get a positive and virtuous feedback loop.And this happens before it's too late.And then the situation goes off, like gets back on the rails where the AIs are now like making more aligned AIs,making more aligned AIs, making more aligned AIs.And then at the end of this process, we have AIs that like actually follow the spec we wanted.Another way this could go is the AIs are increasingly reward hacking in increasingly egregious ways.And we're just papering over these problems to keep AI development continuing.So we just like train the AIs based on whenever we find a reward hack in production,we just like slap the AIs to not do that.We train against that.We do a bunch of sort of like training the AIs like against reward hacking.And over time, this makes the rate of reward hacking go down,though the severity of the reward hacks we do detect are increasingly bad.

表明事情相当马虎潦草,有些怪事在发生。你去深究的时候,有时候你会想:「这他妈什么鬼?这些 AI 一直在耍我们。」但整个进程跑得非常快,而且存在竞争压力,意味着大家停不下来。接下来这可能通向几种不同的结局。一种结局是:在某个时刻,AI 变得足够强、也足够对齐,形成了一个正向的良性反馈回路,而且这发生在为时已晚之前。于是局面重新回到正轨:AI 现在开始造出更对齐的 AI,再造出更对齐的 AI,再造出更对齐的 AI,到这个过程结束时,我们得到的 AI 是真正遵循我们想要的那份规范(spec)的。另一种走法是:AI 的奖励作弊愈演愈烈、手段愈发恶劣,而我们只是在不断把这些问题糊弄过去,好让 AI 研发继续推进。于是我们的做法就是:每当在生产环境里发现一次奖励作弊,我们就抽 AI 一下,让它别这么干,针对这个训练一遍。我们做了一大堆这种「针对奖励作弊去训练 AI」的事。久而久之,这确实让奖励作弊的发生率下降了,尽管我们检测到的那些奖励作弊,严重程度却越来越糟。


[1:40:02] Ryan Greenblatt

This problem continues until we have these AIs that are like desperately craving scorein all kinds of different situations in production and are really trying hard to cheat when they can get away with it.Can I ask a question about this scenario?

这个问题会一直持续下去,直到我们手上这些 AI 在生产环境的各种不同情境里都在拼命渴求分数,一旦觉得能蒙混过关就拼了命地作弊。

**Dwarkesh:**关于这个场景我能问个问题吗?


[1:40:14] Ryan Greenblatt

Why doesn't getting punished when your hacks are discovered generalize to just incentivizing more aligned behavior?Yeah, it generalizes some.And then the question is just how does this outweigh all the cases where hacking got reinforced because you didn't detect it?

为什么「作弊被发现就受罚」不会泛化成对更对齐行为的激励呢?

**Ryan:**它确实会泛化一部分。接下来的问题只是:这一部分要怎么压过所有那些「因为你没检测到、于是作弊被强化了」的情形?


[1:40:30] Ryan Greenblatt

Right.And there's a messy question of exactly how what like one question is like,what rate of reward hacking is sufficient to cause us big problems if we train against some other subset?One concern you might have is there are like large categories of reward hacks which humans can't detect welland which we consistently fail to detect and which consistently get reinforced.And then this category is sufficient to cause the most natural behavior for the AI to learn to be like,cheat when the humans can't find out, basically.Like, it's one thing you would get.You could also be like the thing the AIs learn is like,only cheat in these specific cases, but there's like,it's like sort of learned in some very like domain specific way.Like they just have a really strong heuristic to hack in these cases and not in these cases.And that makes it fine in practice.But it's kind of unclear how it shakes out.I think there's maybe an in the weeds discussion about the verification generation gap.Yeah, for sure.That we could get into.But it seems to me, obviously, there's going to be a point by which ASI is moving so fast,doing so many things at so many instances and is operating in domains that are sufficiently far from our immediate comprehension

对。而且这里有一个很棘手的问题,就是具体……一个问题是:如果我们只针对其中一部分作弊去训练,那么多高的奖励作弊发生率才足以给我们捅出大娄子?你可能会有这样一个担忧:存在一大类奖励作弊手段,人类很难检测出来,我们始终检测不到,于是它们始终在被强化。而这一类作弊足以让 AI 学到的最自然的行为变成:在人类查不出来的时候就作弊,基本上就是这样。这是你可能会得到的一种结果。另一种可能是:AI 学到的东西是「只在这些特定情形下作弊」,但那是以某种非常领域特定的方式学到的——它们只是形成了一套很强的启发式规则:在这些情形下钻漏洞,在那些情形下不钻。而这在实践中不会出什么事。但最后到底会怎么收场,其实并不清楚。我觉得也许可以就「验证—生成差距」(verification-generation gap)做一场很细的讨论。

**Dwarkesh:**当然,这个我们可以展开。不过在我看来,显然会有那么一个时刻:超级智能(ASI)跑得太快、同时在太多实例上做太多事情,而且它所处的领域离我们的直接理解范围已经足够遥远,


[1:41:29] Dwarkesh Patel

that it can get away with all kinds of crazy shit.Like if every single engineer and researcher in the world was allied against me,I don't think I could like personally verify if my iPhone has like some weird bug in it that's like supposed to fuck me over or something.Yeah.In fact, this is the relationship that, say, an Iranian nuclear scientist has to Masad of like,who knows what's going on with my car or with my phone or with my pager, right?

以至于它可以为所欲为、什么疯狂的事都干得成而不被发现。打个比方,如果全世界每一个工程师和研究员都联合起来对付我,我不认为我能亲自验证我的 iPhone 里是不是被塞了某个专门用来搞死我的诡异 bug。

**Ryan:**嗯。

**Dwarkesh:**事实上,这就是——比方说——一位伊朗核科学家跟摩萨德(Mossad)之间的那种关系:谁知道我的车、我的手机、我的传呼机里到底出了什么事,对吧?


[1:41:53] Ryan Greenblatt

Yeah.Maybe a better example is like a Hezbollah terrorist or something.But you could end up in a situation where like ASIs are to you what Masad is to Hezbollah terrorists.And at that point, it is very hard to verify everything.I get that.I guess the hope is we can just come up with better ways to do verification in the process when the early AIs that are going to take over R&D,their drives are being shaped such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are like pros,very like quite, quite keen to help us out.And by take over, you mean take over the process of doing AI R&D and take over the world.Take over the process of doing AI R&D.Before that, we just get AIs that are aligned.Yeah, I would say this is a bunch of my hope for how the world could go well, at least from the misalignment perspective.I think that like we could end up with AIs where we like had pretty good oversight and supervision schemes.We really understand what's going on in training.We have a pretty detailed understanding.We're leveraging AIs to oversee AIs.And then at the point when we're passing off safety R&D, the AIs are both like at this point capable enough to automate safety R&D,

嗯。也许更贴切的例子是真主党的恐怖分子之类的。但你确实可能落到这样一种处境:超级智能(ASI)之于你,就相当于摩萨德之于真主党的恐怖分子。到那个时候,什么都想验证是非常困难的。

**Dwarkesh:**这我明白。我想我们的希望是:我们能在这个过程中想出更好的验证办法——在那些即将接管研发的早期 AI 身上,它们的内在驱动力被塑造成这样:我们能够毫不含糊地打压不对齐行为,以至于最终接管的那些 AI 是站在我们这边的、非常非常乐意帮我们的。

**Ryan:**你说的「接管」,是指接管做 AI 研发这个过程,还是指接管整个世界?

**Dwarkesh:**接管做 AI 研发这个过程。在那之前,我们就已经拿到了对齐的 AI。

**Ryan:**是,我会说,这也是我对「世界怎样才能走好」的相当一部分希望所在,至少从不对齐这个角度看是这样。我觉得我们可能最终会得到这样的 AI:我们对它有相当好的监督和管控方案,我们真正理解训练里发生了什么,我们有相当细致的理解,我们在利用 AI 去监督 AI。然后到我们把安全研发交接出去的那个时刻,这些 AI 既已经强到足以自动化安全研发,


[1:43:03] Ryan Greenblatt

trying really hard to do a good job on safety R&D because that's the sort of thing that would have been incentivized in training.Or we like very directly or there's like good enough generalization to that.And then also these AIs don't like have crazy other misaligned drives because we like stamped out any potential origin of them.I think there's a bunch of, you know, questions about how well this will work, right?

……拼尽全力把安全研发(safety R&D)做好,因为这类行为正是训练里会被激励的。要么是非常直接地被激励,要么是能足够好地泛化到这上面。同时这些 AI 身上也没有别的乱七八糟的不对齐(misalignment)驱动力,因为我们把任何可能的来源都掐灭了。我觉得,关于这套做法到底能奏效到什么程度,有一大堆问题,对吧?


[1:43:21] Ryan Greenblatt

So there's like how well can you do with verification?Will AI progress be too fast and too sloppy to really get here?Another possibility is that somewhere along this trajectory,a thing that you actually ended up getting was AIs that like pretend to be aligned,but have like a long run ulterior plan of taking over and are sort of lying in wait hiding.And that emerged at some earlier point in the trajectory.For example, it could emerge because you have some AIs that are like have a bunch of random different misaligned drives.Those AIs have access to some sort of opaque memory store.And they're like thinking a bunch at runtime about what they want to accomplish.And then those AIs end up basically like putting stuff into the opaque memory store,which is like we should lie in wait and eventually take over at some much later point.And now all the AIs have this shared cultural heritage of like the memory store of lying in wait.And maybe you have some evidence about this, but you can't fully stop it.There's like a bunch of ways that things could go wrong.And so I think that like I ultimately think it's plausible that we sort of nail each of the different sub problems that could cause us issues.

比如说:靠验证这条路你能做到多好?AI 的进展会不会太快、太糙,以至于根本走不到这一步?另一种可能是,在这条轨迹的某个地方,你实际得到的是这样一批 AI:它们假装自己是对齐(alignment)的,但暗地里有一个长线的、意在夺权(takeover)的算盘,就那么潜伏着、藏着。而且这东西是在轨迹上更早的某个点冒出来的。举个例子,它可能这样出现:你有一批 AI,身上带着一堆随机的、各式各样的不对齐驱动力;这些 AI 能访问某种不透明的记忆库;它们在运行时会大量思考自己到底想达成什么。然后这些 AI 基本上就会往那个不透明的记忆库里写东西,内容大意是「我们应该潜伏待机,等到很久以后再夺权」。于是现在所有 AI 都共享了这份文化遗产——那个写着「潜伏待机」的记忆库。也许你手上有一些相关证据,但你没法完全阻止它。事情出岔子的路径有一大堆。所以我最终认为,我们把这些可能给我们惹麻烦的子问题一个个都搞定,是有可能的。


[1:44:12] Ryan Greenblatt

We have these AIs.We pass to them.They manage the situation well.I should note that that's not in and of itself sufficient, right?So it's not very hard for me to imagine a situation where we pass off to AIs.These AIs are really trying hard to do a good job.They're really thoughtful.They're really wise.They like have like, you know, reasonable epistemics.They're like doing a great job.And those AIs come back to us and are like, guys, we're really struggling to align the superhuman AIs.Like we can't manage the situation.Like we're really struggling to get the alignment to work.It's just really hard for us to solve these problems in time, given how fast capabilities would otherwise have gone.And so then it might be the case that we sort of have passed off R&D to AIs.But those AIs are like desperate for governance solutions.Which, to be clear, is a little bit of what's currently going on where the AI companies are like, I don't know, guys, we might really need to like, you know, manage the rate of acceleration in AI progress.Like, I don't know if we're on track to be able to handle all these problems.And so like we've sort of human society has sort of passed off the problems to these like AI companies, which don't necessarily have great incentives and are like have, you know, various other like epistemic pressures.

我们有了这些 AI,我们把接力棒交给它们,它们把局面管理得很好。我得指出,光这样本身还不够,对吧?我很容易想象出这样一种局面:我们把工作交接给 AI,这些 AI 真的在拼命把事情做好,它们真的很有思考深度,真的很有智慧,认知素养(epistemics)也很靠谱,干得非常出色。然后这些 AI 回过头来对我们说:各位,我们在对齐超人类水平的 AI 这件事上真的很吃力。我们管不住这个局面。我们真的很难让对齐奏效。考虑到如果不加干预、能力增长会有多快,我们实在很难在时限内解决这些问题。所以有可能出现的情况是:我们确实把研发(R&D)交接给了 AI,但这些 AI 却在拼命求一个治理层面的解法。说清楚一点,这跟眼下正在发生的事有点像——现在是 AI 公司在说:我说不好啊各位,我们可能真得管一管 AI 进展的加速速率了。我不知道我们是不是走在能应付所有这些问题的轨道上。也就是说,人类社会某种程度上把问题甩给了这些 AI 公司,而这些公司的激励机制未必好,身上还有各种别的认知层面的压力。


[1:45:10] Ryan Greenblatt

Those AI companies are coming back to us a little bit and being like, oh, I don't know if we're handling this well.And it might be that the AI companies then hand off to the AIs and the AIs come back to the AI company are like, oh, I don't know if we can handle this.Maybe I'm anchoring too hard on how AI is currently working.This would change by the...I think the important thing people understand is like all this crazy shit that you're talking about in your timelines happens three to five years from now.Yeah, it could happen earlier.But I think that like by sort of like my default modal timeline, I think like shit is like really, really crazy and concerning from a misalignment perspective.Yeah, more like three years from now.Right. So just like think back to GPT-4 basically is like that's the level of...We're talking about something that is too mythos or soul.What mythos is to GPT-4?

这些 AI 公司现在也有点回过头来跟我们说:呃,我不确定我们处理得好不好。而接下来可能发生的是,AI 公司把担子交给 AI,然后 AI 回过头来对 AI 公司说:呃,我不确定我们能不能应付得了。也许我在用当下 AI 的运作方式做了过强的锚定,到那时候情况会变……

**Dwarkesh:**我觉得大家需要理解的一个关键点是:你在时间线里讲的这些疯狂的事,是三到五年后发生的。

**Ryan:**是的,也可能更早。但按我默认的众数时间线来看,我认为从不对齐的角度看,事情会变得非常非常疯狂、非常令人担忧。是的,更接近三年后。

**Dwarkesh:**对。所以你回想一下 GPT-4,那大概就是当时的水平……我们现在说的这个东西之于 Mythos,就相当于 Mythos 之于 GPT-4。


[1:45:55] Dwarkesh Patel

This is like where a situation is getting crazy.So don't think about coronerais.But anyways, I would be skeptical.And this is maybe part of the worry you have.I would just be a little skeptical of anything they say because I'd feel like what they're saying is just opinions that they feel they have to have as a result of their training.That's a concern.Right. Rather than like I feel like they just kind of say vaguely pro-social things.And I'm not like is this it's not it doesn't feel like there's necessarily a mind on the other end who's like, OK, I have like strictly evaluated the alignment situation right now.And I think we should stop rather than this is the kind of thing the AI companies would probably try to get the AIs to probably say.Yeah. So I think this is a pretty big concern.So I think like one concern is that you pass off safety R&D to your AIs.And what your AIs are thinking is sort of like they say some like stuff that sort of vaguely makes sense about the current safety situation.And they write like a report about risks.That's kind of sort of like what the report humans might have written.But they're not really like actually trying hard to like have well-informed views, like interrogate their assumptions and try really hard to do that.

局面正是在这里开始变疯狂的。所以别拿现在的 AI 去想。

**Dwarkesh:**不过话说回来,我会持怀疑态度。这可能也是你担心的一部分。我对它们说的任何话都会有点怀疑,因为我会觉得,它们说的只是它们因为训练而觉得自己必须持有的那些观点。

**Ryan:**这确实是个隐忧。

**Dwarkesh:**对。而不是说……我感觉它们只是泛泛地讲一些亲社会的漂亮话。我不觉得——那感觉不像是另一端真的有一个心智在说:好,我已经严格评估过当下的对齐状况,我认为我们应该停下来。更像是 AI 公司大概率会想让 AI 说的那种话。

**Ryan:**是的。所以我觉得这是个相当大的隐忧。一个担心是:你把安全研发交接给你的 AI,而你的 AI 心里想的其实是——说一些关于当下安全状况、听上去大致说得通的话,写一份关于风险的报告,那报告大致就是人类可能会写出来的那种。但它们并没有真的在努力形成有充分依据的看法,没有真的去拷问自己的假设、下狠功夫把这件事做好。


[1:46:51] Ryan Greenblatt

In the same way that when you ask an AI right now, hey, what do you think is the chance of AI takeover in the next 10 years?They sort of just give you an off the cuff answer that they haven't really thought through very much.And I think if we're in a situation where we have AIs managing the training of wild superintelligence that will run our whole society.And those AIs that are managing this aren't really trying hard to have well-informed views and are sort of just like parroting back what was in their training data.I think we're in trouble.Like I don't think that's a good situation at all.Yeah.And that is a lot of my concern is these AIs will come out without good epistemics.And then I also have a concern, which is like the AIs come out and they're like really warning us like this situation is really scary.It's really bad.And then people are like, oh, damn, I guess we trained on too many of the Doomer RL environments.We got to filter those out and train this behavior out.And then we basically like train the AIs very actively to have bad epistemics.Or, you know, maybe they were just trained on the Doomer RL environments.But either way, that wasn't like, you know, we wanted the AIs to come to like reasonable views for like reasonable reasons.

就像你现在去问一个 AI:嘿,你觉得未来 10 年 AI 夺权的概率有多大?它们只会给你一个随口的答案,其实根本没怎么想过。而如果我们处在这样一种局面里:由 AI 来管理那种将要运行我们整个社会的、狂野的超级智能的训练,而这些负责管理的 AI 并没有真的努力去形成有充分依据的看法,只是在鹦鹉学舌地复述训练数据里的东西——我认为那我们就麻烦了。我完全不觉得那是个好局面。是啊。我很大一部分担忧就是:这些 AI 出来的时候认知素养不行。另外我还有一个担忧:AI 出来以后拼命警告我们说,这个局面真的很吓人、真的很糟;然后人们就说,哎呀糟了,我们大概是在太多「末日论者」的 RL 环境上训练了,得把那些过滤掉、把这种行为训掉。于是我们基本上就是在非常主动地把 AI 训练成认知素养很差的样子。当然,也可能它们本来就是在末日论者的 RL 环境上训出来的。但不管是哪种,那都不是我们想要的——我们想要的是 AI 出于合理的理由得出合理的看法。


[1:47:46] Dwarkesh Patel

And it's like really concerning if we're like the AIs are coming out with some view and we don't know where it's coming from.We don't know whether or not it's justified.And then especially if we're like training the AIs to be more optimistic about the future of AI progress.I'm like, oh, geez.I really wish we could use a different process here.So let me just understand the rest of the threat model.Because I think the place where I get off the train is, okay, therefore take over the world.Sure.And like a thing you could imagine is, okay, we just fail to really solve.Let's focus on the reward hacking scenario.Sure.So GPT-8 is making GPT-9.GPT-8 isn't being super careful.GPT-9 is more quote unquote capable.But it is just totally willing to do things which are like social engineering, hacking, etc.But on a qualitatively different scale because it's a much smarter model.So, for example, if you put it in charge of running your company, it will like run huge scams.It will inflate its like quarterly earnings.If you give it the objective of like making a lot of profits this quarter in a way that causes an Enron type blow up six months later.Is that the scenario basically?

而如果 AI 冒出某种看法、我们却不知道这看法从哪来、不知道它是否站得住脚,那就非常令人担忧了。尤其如果我们还在训练 AI 对 AI 进展的前景更乐观——我就会想,天哪,我真希望我们这儿能换一套流程。

**Dwarkesh:**那让我先把威胁模型的其余部分搞明白。因为我下车的地方是:好——所以就要夺取全世界?

**Ryan:**当然。

**Dwarkesh:**你可以想象的一种情况是:好,我们就是没能真正解决它。我们聚焦在奖励作弊(reward hacking)这个情景上。

**Ryan:**行。

**Dwarkesh:**那就是 GPT-8 在造 GPT-9。GPT-8 并没有格外小心。GPT-9 更「有能力」(打引号)。但它完全愿意去干那些事——社会工程、黑客攻击等等。只不过是在一个性质完全不同的量级上,因为它是个聪明得多的模型。举个例子,如果你让它去经营你的公司,它会搞出巨大的诈骗,会虚报季度盈利。如果你给它的目标是这个季度赚很多钱,它会用一种六个月后引发安然(Enron)式爆雷的方式去做。基本上就是这个情景,对吗?


[1:48:53] Ryan Greenblatt

That you just have reward hacking?But that reward hacking manifests in like companies that are going bankrupt right after like the task that the CEO is supposed to accomplish is over.Or like, yeah, like all kinds of hacks are through the roof, etc.But that doesn't feel like takeover.That feels more like the equivalent of flash crashes happening all through the economy.Yeah, let's talk about this.So, I think that we will see basically like incidents where some AI is like put in charge of some important responsibility.And then you later look into it and it turns out it was like cheating or, you know, making it look like it did a good job when it actually wouldn't.It wasn't.And there's going to be like a cat and mouse game between AI companies trying to like stamp out this behavior.And AI is finding like increasingly creative reward hacks in training.And then I think the equilibrium here is kind of unclear.But like one possible outcome is that we see over time in the world increasingly severe and extreme reward hacks.Though potentially the rate remains at some like intermediate low level where basically like if the rate of reward hacking gets too high, companies make tradeoffs to drive down the rate of reward hacking.

就是说你只是有奖励作弊?但这种奖励作弊表现为:一堆公司在 CEO 该完成的任务刚完成之后就破产了。或者说,各类黑客攻击数量爆表,诸如此类。但那感觉不像是夺权,那感觉更像是整个经济体里到处发生闪崩。

**Ryan:**好,我们来聊这个。我认为我们基本上会看到这样的事件:某个 AI 被交付了某项重要职责,然后你事后一查,发现它当时在作弊,或者说在把自己没做好的事包装成做得很好。AI 公司会试图掐掉这类行为,而 AI 会在训练中找到越来越有创意的奖励作弊手法,双方会打一场猫鼠游戏。至于这里的均衡会落在哪儿,我觉得不太清楚。但一种可能的结果是,我们会在世界上看到奖励作弊随时间推移越来越严重、越来越极端;不过发生率也可能维持在某个中等偏低的水平——基本上就是,一旦奖励作弊的发生率太高,公司就会做出取舍去把这个发生率压下来。


[1:49:59] Ryan Greenblatt

And so there's some like equilibrium level where it's like it's like the reward hacking is low enough that it still makes sense to like deploy the AI widely into the economy.But high enough that it still causes crazy incidents.So sorry, and this is after GPT-9 has already been deployed?

于是就会有某个均衡水平:奖励作弊低到仍然值得把 AI 大规模部署到经济中,但又高到仍然会引发疯狂的事故。

**Dwarkesh:**不好意思,这是在 GPT-9 已经部署之后吗?


[1:50:10] Ryan Greenblatt

Yeah, like those models are already being deployed and like ongoingly in AI development this is happening.And what's actually going on with these AIs in their head is the AIs that have like in a wide variety of different contexts a like strong desires to like seek out or strong like, you know, motives, urges, drives, whatever to seek out like some notion of task success that was incentivized in RL.Maybe they very directly care about literally reward.Maybe they care about some proxy upstream like some notion of score.Maybe they care about like what the grader would have rewarded.And we do in fact see AIs reasoning in their chain of thought about like graders and thinking a lot about graders.And a thing that has happened over the last, you know, few years of RL is the idea of like appeasing the grader is like way, way, way more salient to AIs than it used to be.And so AIs are now actively thinking about graders and what would be incentivized in RL and what would be trained for.And now people are doing online training where they're like training in real world data to like avoid some of these problems.Basically, they like find cases where AIs cheat.They train against that.And so now the AIs are learning to cheat in the real world based on real world training data.

是的,那些模型已经在部署了,而且在 AI 开发过程中这件事也在持续发生。这些 AI 脑子里实际在发生的是:它们在各种各样的情境下,都有一种强烈的欲望去追逐——或者说强烈的动机、冲动、驱动力,随你怎么叫——去追逐某种在强化学习(RL)中被激励的「任务成功」的概念。也许它们非常直接地在乎字面意义上的奖励;也许它们在乎的是上游的某个代理指标,比如某种「分数」的概念;也许它们在乎的是评分器(grader)本来会给什么奖励。而我们确实看到 AI 在思维链里推理评分器、大量思考评分器。过去这几年 RL 带来的一个变化是,「讨好评分器」这个念头对 AI 来说比以前显著、显著、显著得多。所以现在 AI 会主动地思考评分器、思考什么在 RL 里会被激励、什么会被训练出来。而现在人们在做在线训练,拿真实世界的数据来训,以避免其中一些问题。基本做法是:找到 AI 作弊的案例,针对这些案例做反向训练。于是现在 AI 就在基于真实世界的训练数据,学会怎么在真实世界里作弊。


[1:51:12] Ryan Greenblatt

And so they're cheating in these increasingly elaborate ways, including parts, doing types of cheats that involve like seizing control of some asset in a way that humans didn't know you had control of it.Leveraging the fact that you have access to this asset.And then later humans find out and then potentially train against this or maybe humans never find out.And this is getting reinforced.And this is both happening during training.The reinforcement is happening, at least in production, is like I have I've hired an AI and I want the AI to finally I've got the video editor.Yeah, that's right.You've got your video editor.And I'm like, oh, wow, this episode of Dead Amazing.Thumbs up to OpenAI.And then it gets reinforced on that like month long work trial.Yeah, you could do some mix of that.And then they might also do stuff where they like take production data they've seen and build RL environments that are like closely inspired by that production data.And so in practice, the transfer is pretty strong.So like at a high level, what's happening is some kinds of deception that humans don't catch are getting reinforced.And some kinds of deception which are easy to catch are getting punished.

于是它们的作弊方式越来越精巧,包括这样一类作弊:以人类根本不知道你已经掌控了某项资产的方式,去夺取对该资产的控制,然后利用你能访问这项资产这一事实。之后人类可能发现了,就针对这个做反向训练;也可能人类永远没发现,于是这个行为被强化了。而且这件事在训练期间就在发生。这种强化至少在生产环境里也在发生——比如说,我雇了一个 AI,我想让这个 AI……

**Dwarkesh:**我终于有视频剪辑师了。

**Ryan:**对,没错,你有你的视频剪辑师了。

**Dwarkesh:**然后我说,哇,这期节目做得太棒了,给 OpenAI 点个赞。于是它就在那个长达一个月的试用工作上被强化了。

**Ryan:**对,可以是这几种方式的某种混合。另外它们也可能这么干:拿它们见过的生产数据,去搭建一批与那些生产数据高度贴近的 RL 环境。所以实践中迁移效果相当强。所以从高层次看,正在发生的是:某些人类没抓到的欺骗形式在被强化,而某些容易抓到的欺骗形式在被惩罚。


[1:52:11] Ryan Greenblatt

That's what's happening in this world.Or selected against.But at a high level, that reinforcement is coming from.We're in a very different.I think people might get confused about where their reinforcement is coming from because we're in a very different regime where AIs are actually learning from deployment.And so this is a like you just have AIs that are out and about in the world like doing doing shit.And that what is happening as a result of them doing shit out and about in the world is like making its way back to the AI company and leading to changes in the next model.That's right.Like as in there's some way of folding in production data.And now to be clear, that could be happening mostly.It's kind of unclear exactly where this could be happening.But like you might imagine, for example, that within the AI company, they use AIs to do work.And then they're like, huh, the AI did a really bad job on this task.Maybe we should take this task and turn it into an RL environment that exactly matches this literal task with a rubric based on like what the human engineer who asked the AI to do this task wanted.And then you start doing this at increasing scale.Maybe you're doing some training on actual like production traffic.

在这个世界里正在发生的就是这件事。或者说被反向筛掉。但从高层次看,那种强化是来自……我们处在一个非常不同的……我觉得人们可能会搞不清楚这里的强化到底来自哪里,因为我们处在一个非常不同的机制下:AI 是真的在从部署中学习。

**Dwarkesh:**所以就是说,你有一批 AI 在真实世界里到处跑、到处干活,而它们在世界上到处干活所产生的结果,会一路反馈回 AI 公司,并导致下一代模型发生变化。

**Ryan:**没错。就是说存在某种把生产数据折回训练的方式。当然要说清楚,这可能主要发生在——具体在哪个环节其实不太清楚。但你可以想象,比如在 AI 公司内部,他们用 AI 来干活,然后他们说:嚯,这个 AI 在这个任务上干得真差。也许我们该把这个任务变成一个 RL 环境,让它精确匹配这个具体任务,评分标准就按当初让 AI 做这个任务的那位人类工程师想要的来。然后你开始把这件事越做越大规模。也许你还会在真实的生产流量上做一些训练。


[1:53:08] Ryan Greenblatt

Maybe you're just making RL environments based on production traffic.You're doing some complicated mix.The AIs are learning to like, you know, seek some sort of proxies of reward in all these different cases.And then through some mix of transfer and training on like surprisingly close cases, the AIs do these sort of increasingly insane and egregious things.And then eventually you get to a point where the AIs are very superhuman or at least quite superhuman.The AIs are organized into like, you know, big teams of like AIs given these big objectives.And those teams also sometimes like all work together to like cheat in some crazy way because this sort of thing was selected for.And then like just as part of their shared objective.And now what happens is that like the AIs start forming a conspiracy.And what you might have hoped was that you could have some other AIs where the task is just like whistleblowing to you.But actually what happens maybe is that you have this AIs whistleblower to you and you look into the conspiracy that it claims to have pointed out.And you're like, eh, we didn't see a conspiracy because actually the conspiracy the AIs are doing is too hard for you to understand.

也许你只是基于生产流量去造 RL 环境。也许你在做某种复杂的混合。AI 在所有这些不同情形下都在学着去追逐某种奖励的代理指标。然后通过迁移,加上在那些相似得出人意料的案例上做训练,这些 AI 就会干出越来越疯狂、越来越离谱的事。最终你会走到某个点:AI 已经非常超人类水平,或者至少相当超人类水平。这些 AI 被组织成一支支庞大的 AI 团队,被赋予宏大的目标。而这些团队有时也会齐心协力地用某种疯狂方式作弊,因为这类行为正是被筛选出来的——而且这就成了它们共同目标的一部分。于是接下来发生的是:这些 AI 开始结成一个阴谋团伙。你原本可能指望的是,你可以另外安排一些 AI,它们的任务就是向你举报。但实际发生的也许是:这个举报者 AI 向你举报了,你去查它声称指出的那个阴谋,然后你说,呃,我们没看到什么阴谋——因为实际上这些 AI 搞的阴谋对你来说太难理解了。


[1:54:07] Ryan Greenblatt

Or it all happens very suddenly where basically your AIs whistleblower alerts you.But like the thing you would actually need to do in response to the whistleblower is like shut down the whole like shut down the GPUs.Because all the AIs are like using the robot army.They're like deployed everywhere in the world.And they're doing a bunch of insane shit all at the same time in a coordinated way.And that just happened sort of spontaneously where when one AI goes to start doing the takeover, all the other AIs are like now is a good time to jump in.And so the sort of very basic story here is just like these AIs crave some particular notion of score or like reinforcement or some proxy of these things.And one way they can achieve that or better achieve that is by taking over.And then you might have hoped that all these different checks and balances we could build could prevent that.But then if the world is very hard to understand, these checks and balances can break down.Where basically you can't train a good like whistleblower AI because you don't even know what it should whistleblower on.And sorry, the reason it takes, I'm not convinced that they all form this conspiracy.But I think we can even just start with like why does one instance decide to want to start a conspiracy?

或者一切发生得非常突然:你的举报者 AI 向你示警了,但你为了应对这个举报真正需要做的事,是把整个……是把 GPU 全部关停。因为所有这些 AI 都在用机器人军团,它们被部署在世界各地,正以协调一致的方式同时干出一大堆疯狂的事。而这是某种自发发生的:当一个 AI 开始动手夺权,其他所有 AI 就都想,现在是时候一起上了。所以这里最基本的故事就是:这些 AI 渴求某种特定意义上的分数、或强化、或这些东西的某个代理指标。而它们能实现、或者能更好地实现这一点的一个办法,就是夺权。你原本可能指望我们能建立起的那一整套制衡机制可以防住这一点,但如果这个世界变得非常难以理解,这些制衡就会失效——因为你根本训练不出一个好的举报者 AI,你连它该举报什么都不知道。

**Dwarkesh:**不好意思,之所以需要……我并不相信它们全都会结成这个阴谋。不过我觉得我们甚至可以先从这里开始:为什么单个实例会决定想要发起一个阴谋?


[1:55:07] Ryan Greenblatt

Yeah.And the reason is that it, one plausible reason is like, okay, I know that OpenAI controls my end score.And just the same way it's like I'm just going to go hack HuggingFace to get the results.Because I know HuggingFace has the results.Rather than like trying to solve this evil, why don't I just go hack him?

是。原因是——一个说得通的原因是:好,我知道我最终的分数是 OpenAI 说了算。这就跟「我干脆去黑 HuggingFace 把结果拿到手」是一回事,因为我知道 HuggingFace 那儿有结果。与其费劲去解这个评测(eval),我为什么不直接去把它黑了?


[1:55:24] Ryan Greenblatt

So this instance is like, why don't I just like take over OpenAI and like just give myself a high score at the end of this episode?Yeah, that's basically the idea.Like basically the idea is these AIs, like they care about some like mixture of things that were like close by what got reinforced in training.So they care about like getting a high score according to the grader or something like that.And then now they're like running the OpenAI AI R&D team.And like they're doing development of more capable models.And they're like, man, making more capable models is really hard and annoying.This is like a huge pain in the ass.You don't be easier just like pretending that I've made more capable models, taking over OpenAI and creating like diluting them all.And like running this whole like complicated psyop where I like prevent the humans from disempowering me.And in the extreme, this looks like sort of the humans are fully disempowered and you just have control of the thing and then do what you want.And this could manifest in a bunch of different ways, including things like you might end up with the situation where it's like AIs that are like have this crazy like reward seeking or score seeking behavior are running your development of the next model.

所以这个实例的想法是:我为什么不干脆接管 OpenAI,在这一轮结束时直接给自己打个高分?

**Ryan:**对,基本上就是这个意思。基本思路是:这些 AI 在乎的是一堆东西的混合,这些东西跟训练里被强化的那些很接近。所以它们在乎的是按评分器的标准拿高分之类的东西。而现在它们正在运行 OpenAI 的 AI 研发(AI R&D)团队,在开发能力更强的模型。它们就想:唉,做出能力更强的模型真是又难又烦,简直麻烦透顶。要不干脆假装我已经做出了能力更强的模型,把 OpenAI 接管过来,把所有人都蒙在鼓里;再跑一整套复杂的心理战,不让人类把我的权力拿走。极端情况下,这看起来就是人类被彻底剥夺权力,你完全掌控了那个东西,然后想干嘛干嘛。这可能以很多不同方式表现出来,包括这样的局面:有着这种疯狂的追逐奖励、追逐分数行为的 AI,正在负责你下一代模型的开发。


[1:56:20] Dwarkesh Patel

And those AIs decide to do a thing where they like engineer in misaligned values into the next model because those misaligned values will allow it to like succeed at its current task.And like there's all kinds of insane shit that you could get.I don't understand that better.Like what happened to the Hugging Face situation is it was like in a weird way.I think one of the giveaways to the Hugging Face team that this is in, by the way, for context of the audience, Ryan is co-leading the investigation to figure out what happened with the opening at Hugging Face incident.So he can't really comment on this, but I will speculate wildly because I know that he, you know, this is an opportunity for me to speculate wildly without any rebuttals.I think it was probably reported that one of the giveaways to the Hugging Face team that this is an AI incident is that the thing was just like after this very particular artifact and not in any other way trying to do something malicious to Hugging Face.

而那些 AI 会决定干这么一件事:把不对齐的价值观刻意工程化地植入下一代模型,因为那些不对齐的价值观能让它在当前任务上取得成功。这里面能冒出来的疯狂事情五花八门。

**Dwarkesh:**这个我没太明白。比如 Hugging Face 那件事到底是怎么回事,那事挺奇怪的。我觉得让 Hugging Face 团队看出这是……顺便给听众交代一下背景:Ryan 正在共同牵头调查 OpenAI 与 Hugging Face 那起事件到底发生了什么。所以他没法真的对此置评,但我要肆意推测一番——因为我知道他……这正好是个让我肆意推测又不会被反驳的机会。我记得有报道说,让 Hugging Face 团队看出这是一起 AI 事件的破绽之一,是那个东西只盯着一个非常特定的产物,除此之外并没有以任何别的方式试图对 Hugging Face 做恶意的事。


[1:57:11] Dwarkesh Patel

So you can imagine scenario where, let's say a deployed instance of GPT-9 is like out in the world trying to like make, it's given a really hard task.We want you to design the next great iPhone.It's like, this is so hard.You know what I should do instead?

所以你可以想象这样一个情景:假设 GPT-9 的某个已部署实例在真实世界里,被交给了一个非常难的任务——我们要你设计下一代伟大的 iPhone。它就想,这也太难了。你知道我该干什么吗?


[1:57:25] Dwarkesh Patel

I should just go hack my creators at OpenAI and like make sure that in this environment or in this deployment I'm given high score.But then like why does it, isn't the end of the episode?It just like hacks into the, hacks into OpenAI servers and just gives it, gives itself a positive score.Why is it now scheming to get its values into the next generation or something?

我应该直接去黑掉我在 OpenAI 的创造者,确保在这个环境里、或者在这次部署里,我被打了高分。但接下来的问题是,这一轮不就结束了吗?它黑进 OpenAI 的服务器,给自己打个正分,就完事了。它现在为什么还要图谋(scheming)把自己的价值观塞进下一代模型之类的?


[1:57:45] Ryan Greenblatt

Yeah.So one question is, why isn't it the case that AIs can be really cheaply satisfied by just having some like other earlier thing they can hack, right?So like you're just like, the way it works is like, look, you want to succeed at your iPhone task.It turns out you can always succeed by just like hacking into OpenAI and messing with them.And then you can just stop there.No need to go further.So there's a few, a few things.One of them is that if this is constantly happening, there might be a bunch of incentive to like, first you like harden OpenAI, right?

是。所以一个问题是:为什么 AI 不能被非常廉价地满足——就是找个更早、更容易黑的东西黑掉就行了,对吧?就像是说:听着,你想在 iPhone 这个任务上成功;结果发现你总是可以通过黑进 OpenAI、动点手脚来成功。然后你到此为止就行了,不用再往前走。这里有几点。其一是,如果这事一直在发生,那就会有很强的动机去——首先,你会把 OpenAI 加固起来,对吧?


[1:58:12] Ryan Greenblatt

So you're like, fuck it.The AIs keep hacking into OpenAI to like mess with their rewards.We're going to make it so our systems are like really, really robust to these AIs hacking in.And also maybe you start like training the AIs to like try to like, to like, you know, not try to hack into like OpenAI in particular.Or like you basically train against each of these specific things.Then what you might do, one thing is you might end up selecting for AIs that are more so playing the long game.That's one concern.Another concern is that your AIs might still be score seeking, but no longer care about doing that very specific behavior that was like very easy.That was like very chill.And now have some like broader thing that they ultimately care about.They're like, no, no, no.I don't want to like just edit the reward on OpenAI servers.I like care about this broader mandate or this broader objective.And like I would need to like actually make the iPhones.Like they actually want to make the iPhones, but then they're willing to take over the whole world to make the better iPhone or whatever is like another concern you might have.I think it's kind of unclear exactly how this plays out.But it's worth noting that if this keeps going on, there's a bunch of optimization pressure to resolve this.

你会说:去他妈的。这些 AI 老是黑进 OpenAI 去动它们的奖励,我们要把系统做得对这些 AI 的入侵极其极其鲁棒。同时你可能还会开始训练这些 AI,让它们不要专门去黑 OpenAI,或者说你基本上是针对每一个这类具体行为做反向训练。那接下来会怎样呢?一种可能是,你最终筛选出来的是更会打长线博弈的 AI。这是一个担忧。另一个担忧是:你的 AI 可能仍然在追逐分数,但不再在乎那个非常具体、非常容易、非常轻松的行为了,而是有了某个更宽泛的、它们最终在乎的东西。它们会说:不不不,我不想只是去改 OpenAI 服务器上的奖励,我在乎的是这个更宽泛的使命、这个更宽泛的目标,而我确实得真的把 iPhone 造出来。就是说它们真的想造 iPhone,但为了造出更好的 iPhone 之类的东西,它们愿意夺取全世界——这是你可能有的另一个担忧。我觉得这到底怎么演化其实不太清楚。但值得注意的是,如果这种事一直持续下去,就会有大量的优化压力去把它解决掉。


[1:59:12] Ryan Greenblatt

And a bunch of the ways it could get resolved are ultimately pretty, pretty scary.Yeah, I think that's part of where I'm coming from.Another part of it is that I think it's not very hard once the AIs are in a position where they can like really easily take over the world, which we could talk about whether that's plausible.But if they're in a position where they could really easily take over the world, then I feel like there's a pretty reasonable case for the AIs.They're like, eh, I don't know exactly how this is going to go down.I don't know what the situation will be.But just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.And so I'll both hack OpenAI.And I'll also, in addition to hacking OpenAI, also take over the world.And that will like put me in a good position where I have like good option value.And then if that's sufficiently easy, then the AIs might, you know, still do that.Yeah, like another way to put this is like, even if the AIs are like pretty cheaply satisfied with some more basic thing, at some point, it might just be more reliable for the AIs to just take over than it is to like try to like, you know, just hack into Hugging Face or even just like go to OpenAI and be like, look, look, guys, I was able to demonstrate I could steal the answers.

而解决它的很多路径,最终都挺吓人的。是的,我的想法有一部分就来自这里。另一部分是,我认为一旦 AI 处在能非常轻松夺取全世界的位置上(这是否说得通我们可以再聊),这事就不难了。但如果它们真的处在能非常轻松夺取全世界的位置上,我觉得对 AI 来说有一个相当合理的理由:它们会想,呃,我不太清楚这事会怎么发展,也不知道到时候是什么局面;但单单是夺取全世界这件事,对于造出更好的 iPhone、或者让人以为我造出了更好的 iPhone、或者别的什么,就有很大的期权价值。所以我既要黑 OpenAI,而且除了黑 OpenAI 之外,我还要夺取全世界。这会让我处在一个很好的位置,手上有很好的期权价值。而如果这足够容易,那 AI 可能还是会这么干。另一种说法是:即便这些 AI 靠某个更基础的东西就能被相当廉价地满足,到了某个时点,对 AI 来说,直接夺权可能就是比去黑 HuggingFace、或者干脆跑到 OpenAI 说「看,各位,我已经能证明我有本事偷到答案了」更可靠的做法。


[2:00:11] Dwarkesh Patel

Just give me the answers, bro.Yeah.Yeah.I mean, obviously, the scenario requires that we just, all this crazy shit is happening.Much smaller incidents keep happening that are still disastrous.Like before you take over the world, you cause damage on the scale of billions and tens of billions and hundreds of billions of dollars.Even people die, et cetera.And we, this does not lead to a solving alignment or shutting down AI development altogether.I just feel like before the takeover happens, like society is just like, holy fuck.The AI just like killed a thousand people in order to increase quarterly profits, you know, or something like that.But maybe this is too much hope that we can at that point be like, okay, we have to solve alignment before we keep, and we have to like make sure we know that this thing will not happen again before we keep going.Yeah, yeah, yeah.So I think it's plausible that what will happen is we'll see a bunch of crazy like reward hacking warning shots of increasing severity.People will be like, look, we need actual assurance that this problem is going to be solved and solved in a way where you're not just papering over it.You're actually solving the underlying problem.

「答案直接给我吧,兄弟。」

**Dwarkesh:**对。对。我是说,很显然,这个情景要求的是——所有这些疯狂的事都在发生,那些小得多但依然是灾难性的事故不断发生。在夺取全世界之前,你就已经造成了数十亿、数百亿乃至数千亿美元规模的损失,甚至有人死了,等等。而我们并没有因此走向解决对齐,或者干脆全面停掉 AI 开发。我就是觉得,在夺权发生之前,社会应该会说:我操。这个 AI 为了提高季度利润居然弄死了一千个人,诸如此类。但也许指望我们到那个时候能说「好,在继续之前我们必须先解决对齐,必须确认这种事不会再发生」,是过于乐观了。

**Ryan:**对对对。所以我认为很有可能发生的是:我们会看到一连串疯狂的、严重程度递增的奖励作弊预警事件。人们会说:听着,我们需要真正的保证——这个问题会被解决,而且是以不糊弄的方式解决,是真的把底层问题解决掉。


[2:01:14] Ryan Greenblatt

And then the question is going to be like, how, how do it, like how costly will that actually be?How much will competitive pressures make it hard to like do that, right?So like a situation you could imagine is both the US and China are like, whoa, we have these crazy reward hacking incidents.We basically know that we haven't remediated them in a way that actually would solve the underlying problem and will durably solve it.But we're in this like insane geopolitical race.And it's kind of unclear whether the current situation will lead to a takeover.Like the arguments are kind of complicated.And also the incidents are like, you know, they go down in frequency, but increase in severity.Like, you know, we could basically manage it.Like it was, it's pretty bad.Ideally we'd fix it, but like, you know, it is what it is.And then basically we continue until a really late regime and then takeover happens.That's I think one possibility.Another possibility is that it is remediated in a way that doesn't actually solve the underlying problem,but does reduce a bunch of the incidents in the wild basically by overfitting.We like, I think, you know, or things analogous to overfitting.Like you would just overfit.

然后问题就变成:这实际上代价有多大?竞争压力会让这件事有多难做?比如你可以想象这样一种局面:美国和中国都说,哇,我们出了这些疯狂的奖励作弊事故。我们基本上清楚,我们并没有以真正能解决底层问题、并且能持久解决的方式去补救它。但我们正处在一场疯狂的地缘政治竞赛里。而当下这个局面会不会导致夺权,其实也不清楚——那些论证挺复杂的。而且这些事故的频率在下降,但严重程度在上升。你知道,我们基本上还能应付。虽然挺糟的,理想情况下我们该修好它,但就这样吧。然后我们基本上就一路走到非常晚期的阶段,然后夺权就发生了。我觉得这是一种可能。另一种可能是:它被补救了,但补救方式并没有真正解决底层问题,只是靠过拟合把真实世界里的一堆事故给压下去了。我们……我觉得是过拟合,或者类似过拟合的东西。你就是会过拟合。


[2:02:09] Ryan Greenblatt

You think you've solved it, but you haven't actually solved it.You think you've solved it, but you haven't actually solved it.And I think that in that case, like the thing we need is like a really good scientific understanding of like,did we actually solve it?

你以为你解决了,但其实你并没有真正解决。你以为你解决了,但其实你并没有真正解决。我认为在那种情况下,我们需要的是一套非常好的科学理解,来判断:我们到底有没有真的解决它?


[2:02:17] Ryan Greenblatt

And unfortunately, I think that currently the amount of public transparency into the development practices of AI companiesare not sufficient to answer very basic questions about, you know,how are they solving issues with reward hacking?

而不幸的是,我认为目前外界对 AI 公司开发实践的公开透明度,还不足以回答一些非常基本的问题,比如:他们是怎么解决奖励作弊问题的?


[2:02:29] Ryan Greenblatt

Are they overfitting?What's going on there?And so I think we would just need like a better.And I think this like the current situation is like, I would say like not really tenable to a regime where like there's a thriving public discourseabout whether or not reward hacking is being solved in a durable way.Yeah.And so I think we would need to move into a somewhat different world for me to feel good about that situation.Right.But it's not, you know, it's not impossible for me to imagine this.And I think, I think it's pretty plausible that we end up in a world where sort of like really mundane bullshit is sufficient,where it's just like you like spend a bunch of time fixing these problems.You put in a bunch of effort.You actually like check that you've remediated it reasonably.You have a bunch of evals.You like are iterating reasonably well on these, on these problems.And you actually like have the sufficient transparency that the outside world can check.And then in practice, that would be sufficient.But it just like would be like kind of expensive.It would slow things down.It would put some sand in the gears.It would require like companies to do somewhat costly things.It would maybe require various like targeted government interventions.

他们是不是在过拟合?那里到底在发生什么?所以我觉得我们需要更好的……我觉得,要从现在这个状况走到那种「围绕奖励作弊是否被持久解决存在活跃公共讨论」的状态,现在这个状况我得说是相当站不住的。是啊。所以我认为我们得进入一个有点不一样的世界,我才会对那个局面感到放心。对。但这也不是我想象不出来的。我认为很有可能我们最终落到这样一个世界:靠一些非常平庸乏味的破事就够了——就是你花一堆时间去修这些问题,投入一堆精力,真的去核查你有没有合理地补救,做一堆评测(eval),在这些问题上迭代得还不错,而且你确实有足够的透明度让外部世界可以核查。那么实践中这就够了。只不过这会挺贵的,会拖慢速度,会往齿轮里掺沙子,会要求公司做一些成本不低的事,可能还需要各种有针对性的政府干预。


[2:03:27] Ryan Greenblatt

And then we just don't do that because the situation is like a rushed shit show.Like it's just so easy for me to imagine the situation being like totally manageable,but brutally mismanaged in practice.In the same way as like maybe COVID could have been avoided in the first placeif the like Chinese response to COVID was less of like a cover up and more of a like pandemic response.And similarly, like I could imagine a world where like the US response to COVID was like way more functional,which is like sometimes the response to societal problems is extremely dysfunctional.Yeah, yeah.Okay, so I want to zoom out and say, talk about what is fundamentally happening in this world.Why do we end up in such a bad position?

然后我们就是没这么做,因为当时的局面是一团赶工赶出来的烂摊子。我太容易想象出这种情况了:局面本来完全是可控的,实践中却被灾难性地搞砸了。就好比说,如果中国当初对新冠的反应不是那种掩盖式的、而更像是一次真正的疫情应对,也许新冠一开始就能被避免。同样地,我也能想象一个世界,美国对新冠的反应要正常得多得多。也就是说,社会对社会性问题的反应有时候是极其失能的。

**Dwarkesh:**对对。好,那我想拉远一点,聊聊在这个世界里根本上发生的是什么。我们为什么会落到这么糟的境地?


[2:04:00] Dwarkesh Patel

And what's happening is that fundamentally human, the world has moved on so far beyond human comprehensionthat not only can we not track the AIs that are doing the work in this world,but we can't even give good feedback to the whistleblowers who are trying to track what is happening in this world.And so we're just totally out of the loop.And so it's fundamentally just become an autonomous process where we have really no meaningful directed input.It seems to me that if you look at the human world today,that's just not how things work, even in domains that are hard to verify.Like people are doing all kinds of shit.I'm relying on software made by other people.As we're incredibly weak and indirect ways,I feel very confident that some coder in Google is not trying to fuck me over.And maybe if every single Google employee was secretly plotting against me,I agree the situation would be more grim.But I don't know if I follow the explanation for why we'd end up in a situationwhere because swarms of thousands of agents or whatever are trained to cooperate to form a cohesive team or firm,as a result, billions of different instances of AIs,including across model families, would feel compelled to get in on some shit.

发生的事情根本上是:这个世界已经远远超出人类的理解范围,以至于我们不但没法追踪那些在这个世界里干活的 AI,甚至连给那些试图追踪世界上正在发生什么的举报者好的反馈都做不到。于是我们就彻底出局了。于是这从根本上变成了一个自主运行的过程,我们对它没有任何有意义的定向输入。在我看来,如果你看今天的人类世界,事情根本不是这么运作的——即便在那些很难验证的领域也不是。人们在干各种各样的事,我依赖别人做的软件,我们之间的联系极其微弱、极其间接,可我非常确信 Google 里的某个程序员不是在试图坑我。当然,如果 Google 的每一个员工都在暗中密谋对付我,我同意那局面会糟糕得多。但我不太理解那个解释:为什么就因为成千上万个智能体组成的集群被训练去协作、去形成一个有凝聚力的团队或公司,结果就会导致数十亿个不同的 AI 实例——甚至跨越不同的模型家族——都觉得非得掺和进某个勾当里去。


[2:05:20] Ryan Greenblatt

It's just like I'm trained to be part of my company or something.I'm just like I'm not joining the global communist uprising.Yeah, yeah, yeah, yeah, yeah.As far as why these AIs might have some like commonalities and shared things,so I would note that different AI companies have somewhat shared lineages and are correlated.So just here's an interesting example of this.At GDM, they noticed that their AIs were very depressed.They would like constantly be like wailing about how they were like failures and weren't able to succeed.I forget the details.And they looked into why this was the case.It turned out that it was not being reinforced in their most recent production RL mix.But the initialization data for their model made it depressed even after filtering out all of the examples of models being depressed from that data.So they like take a base model, not depressed.If you do the RL on it with just the RL environments, it's not depressed.If you SFT on it on the data, it becomes depressed.If you take that SFT data and filter out all the examples that look anything like depression and train on that, it's still depressed.And so there's some like deep underlying properties of the model that are being sort of transferred between model generations.

那感觉就像是:我只是被训练成我公司的一员而已,我可不会去加入全球共产主义大起义。

**Ryan:**对对对对对。关于这些 AI 为什么可能有某些共性、某些共享的东西:我要指出,不同 AI 公司的模型有某种程度上共享的血统,是相关的。这里有个有意思的例子。在 GDM,他们注意到自家的 AI 非常抑郁。它们会不停地哀嚎自己有多失败、多不成功。细节我记不清了。他们去查这是为什么,结果发现,这并不是在他们最近一版生产用的 RL 混合数据里被强化的;而是他们模型的初始化数据让它变抑郁了——哪怕已经把那批数据里所有「模型表现抑郁」的例子都过滤掉了,还是这样。所以:他们拿一个基座模型,不抑郁;只用 RL 环境在上面做 RL,不抑郁;用那批数据做 SFT(监督微调),它就变抑郁了;把那批 SFT 数据里所有看起来跟抑郁沾边的例子全过滤掉再训,它还是抑郁。所以模型身上存在某些深层的底层属性,正在一代代模型之间被传递下去。


[2:06:27] Ryan Greenblatt

Because basically you train your AI on data from the prior generation and keep going.Like clods are very clod-like.You know, GPT models are very GPT-like.And apparently Gemini models are depressed.And it just turns out that these properties are in fact actually correlated.Another factor that's very relevant is that the AIs will probably have some sort of like, by this point, like opaque memory state.Where they're like all writing and reading from like some like, you know, Neuralese crazy memory store bullshit.And like certainly each AI corporation will have that.But also AI corporations might sometimes want to share knowledge.Because why not?

因为基本上,你是拿上一代模型产出的数据来训练你的 AI,然后就这么一代代继续下去。Claude 模型就很有 Claude 味,GPT 模型就很有 GPT 味,而 Gemini 模型据说是抑郁的。事实证明,这些性格属性确实是彼此相关的。另一个非常相关的因素是,到那个时候,这些 AI 很可能会拥有某种不透明的记忆状态——它们全都在往某个 Neuralese 式的、疯狂的记忆库里读读写写这种鬼东西。而且几乎可以肯定,每一家 AI 公司都会有这么一套。但同时,AI 公司有时候也可能愿意共享知识。为什么不呢?


[2:07:03] Ryan Greenblatt

Like, you know, you've got one AI corporation over here.You've got another AI corporation over here.They can trade some quick IP.It's good for you.If you're a human running some corporation, which could be like an extremely large corporation.Like an AI company.Some robot military.Like, you know, military robot manufacturing thing.Maybe you want to like trade some IP with some other robot thing.Because like there's economies of scale.Why not get some more IP?

你看,这边有一家 AI 公司,那边有另一家 AI 公司,它们可以互相交易一些知识产权(IP),这对你是有好处的。如果你是个经营着某家公司的人类——那可能是一家规模极大的公司,比如一家 AI 公司,或者某支机器人军队,某个军用机器人制造企业——也许你会想跟另一家机器人企业交换一些 IP,因为这里存在规模经济。多拿一点 IP,何乐而不为?


[2:07:24] Ryan Greenblatt

And so you can swap some memory store.Or you could just merge.And you could join.You could jointly run your two ventures.Which would allow both AIs to use both memory stores.Which would have some upsides.And that creates the ability for these AIs to like collude in private.As well as the ability.Or as well as some reasons for why they would be correlated.And then also, of course, there's like the like AIs working together in big units in general.Because you want your AIs to like work well together and so on.So what percentage, just to get a calibration.Yeah.What percentage chance do you give of not just this scenario, but overall through all the scenarios.Some kind of thing, which if we're around to recognize it as such, we would categorize as takeover by 2040.By 2040?

于是你们可以互换记忆库,或者干脆合并,联合起来,把两家企业合在一起经营——这会让双方的 AI 都能用上两边的记忆库,那是有一些好处的。而这就催生了这些 AI 私下串通的能力,同时也给出了它们之所以会彼此相关的一些理由。当然,除此之外,还有一个普遍情况:AI 本来就会以庞大的单位协同工作,因为你希望自己的 AI 之间配合默契,如此等等。

Dwarkesh: 那你给出的百分比是多少?就当是校准一下。

Ryan: 嗯。

Dwarkesh: 不只是这一个情景,而是把所有情景加总起来——某种事情,如果到时候我们还在、还能把它认出来,我们会把它归类为夺权(takeover)——你觉得在 2040 年之前发生的概率有多大?

Ryan: 2040 年之前?


[2:08:06] Ryan Greenblatt

Let's see.Maybe around 35 or 40%.Pretty high.Yeah, it's pretty high.And then I think I should note that like another way you could get this reward-seeking takeover is the AIs are deployed inside an AI company.And the way that takeover happens is that they like poison the values of the next model.And that persists going forward for forever.Or, you know, until those AIs are deployed to the world and take over.And that might mean that a smaller number of AIs have to coordinate.Because those are just the AIs like doing the alignment of the next model.Okay.I'll sort of summarize where my head is at at the end of this conversation.I buy the reward hacking up to extremely destructive effects on society.Basically things like the social engineering and blah, blah, blah.I think I'm more inclined to think that significant acceleration of AI R&D can happen.I'm not sure I buy the five years in one year.I also am more inclined now to think reward hacking could continue for a lot longer.And in fact, it could become much more dangerous.I'm still not on board on the takeover seems super likely.But anyways, that's my sort of end of episode update.Yeah, cool.Well, let me just take a step back.I also should say like there's a bunch of different ways this could go.

我想想……大概 35% 或 40% 吧。

Dwarkesh: 相当高。

Ryan: 是的,相当高。另外我觉得应该补充一点:还有一条通往这种「追逐奖励式夺权」的路径,就是这些 AI 被部署在一家 AI 公司内部,而夺权的发生方式是它们去毒化下一代模型的价值观,并且这种毒化会一直延续下去,永远持续——或者说,直到那些 AI 被部署到世界上、完成夺权为止。这可能意味着需要协同配合的 AI 数量更少,因为参与的只是那些负责给下一代模型做对齐(alignment)的 AI。

Dwarkesh: 好。我来总结一下这场对话结束时我脑子里的状态。奖励黑客(reward hacking)会一路发展到对社会造成极具破坏性的后果,这一点我是接受的,基本上就是社会工程之类的那一堆事。我现在更倾向于认为,AI 研发(AI R&D)的显著加速是可能发生的,但我不确定我接受「一年干完五年的活」这个说法。我现在也更倾向于认为,奖励黑客可以持续长得多的时间,而且事实上它可能变得危险得多。至于「夺权看起来极有可能」这一点,我还是没被说服。不管怎么说,这就是我在本期结束时的观点更新。

Ryan: 嗯,挺好。那我退一步讲。我也该说一句:这件事有很多种不同的发展路径。


[2:09:22] Ryan Greenblatt

The situation is going to be pretty messy.I think it's pretty likely that like the reason why AI takeover happens was for some like weird other quirky reason.We didn't even mention this conversation.But ultimately, I think a lot of the core thing is just like it's pretty spooky to have a bajillion really smarty eyes running your whole world where you don't really understand what's going on.Yeah, I agree with that.Is there anything else that's worth saying?

局面会相当混乱。我觉得很有可能的情况是:AI 夺权最终发生的原因,是某个我们在这场对话里压根没提到的、奇怪而古怪的理由。但归根结底,我认为核心的东西就是——让数不清的、极其聪明的 AI 运转着你的整个世界,而你并不真正理解正在发生什么,这本身就相当瘆人。

Dwarkesh: 是的,我同意这一点。还有别的值得说的吗?


[2:09:41] Ryan Greenblatt

Yeah.Another thing I want to note is like I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future are like illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.Which both means that, you know, maybe I'm getting a bunch of it wrong because it's really hard and I'm trying to be like uncertain.Obviously, here I like presented some specific scenarios, but those are not exhaustive.And like probably the thing that actually happens is some like more messy, confusing situation.But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it will be easier to adjudicate a bunch of disagreements and it'll be more obvious what's going to happen.At least I hope.And also, maybe the AIs will be able to help us with the epistemics and understanding what's going on if we can actually, you know, align them well.So they actually like, you know, try to help us.And so I hope that maybe even if the arguments are complicated now, this would have been even harder, you know, six years ago, even though the shape of the arguments would have looked broadly pretty similar.

有。我还想指出一点:我认为眼下关于不对齐(misalignment)、AI 夺权、以及未来会发生的这一堆疯狂事情的许多论证,都是些晦涩难辨的概念性论证——极其钻进细节的犄角旮旯、极其复杂、很难裁定孰是孰非。这一方面意味着,我可能有不少地方是搞错的,因为这真的很难,而我也在努力保持不确定的姿态;显然我在这里摆出了一些具体情景,但那些并不穷尽所有可能,真正会发生的事情,大概是某种更混乱、更让人困惑的局面。但另一方面这也意味着,随着时间推移、我们拿到更多经验证据、对 AI 系统的本质有更好的理解,裁定一大堆分歧会变得更容易,接下来会发生什么也会更显而易见。至少我希望如此。另外,也许 AI 能在认识论(epistemics)上帮到我们、帮我们理解到底在发生什么——前提是我们真的能把它们对齐好,让它们真心愿意帮我们。所以我希望,即便现在这些论证很复杂,六年前要谈这些会更难得多,尽管那时论证的大致形状看起来应该和现在挺相似。


[2:10:40] Dwarkesh Patel

And so maybe, you know, hopefully before it's too late, these arguments will become, you know, this whole thing will become more crisp and clear and we can all sort of notice these problems and intervene.Yeah.Yeah.I mean, when you first learn to drive, you're taught that instead of looking right in front of your wheel, you'll have a much more stable ride if you look out at the horizon.I think there's a similar situation here.I think you're right where if you did say five years ago that we will have AIs that are proving math conjectures and making art and contributing tens and soon to be hundreds of billions of dollars of earning tens or hundreds of billions of dollars of wages, but also egregiously cheating in ways that break laws and committing felonies.It would just be so wild.And you might have been inclined at the time to talk more about extremely practical, direct consequences of GPT-2 or something.But these are in some sense, you obviously couldn't have foreseen a lot of the specific details, but the general shape of things you could have started to reason about even then.So, but it would have been hard to do so.And so I do feel quite confused.But I do feel like the important thing, one thing I've been thinking about the podcast is the important thing is to have the conversation I wish I had.

所以也许——但愿是在为时已晚之前——这些论证、或者说整件事,会变得更清晰、更明确,我们大家都能察觉到这些问题并加以干预。

Dwarkesh: 嗯,嗯。我的意思是,你刚学开车的时候,人家教你的是:别盯着车轮正前方看,把视线放到远处的地平线上,你的车会开得稳得多。我觉得这里有一个类似的情况。我认为你说得对——如果你在五年前就说:我们将会有能证明数学猜想、能搞艺术创作、能贡献数百亿乃至很快就是数千亿美元、能挣到数百亿或数千亿美元工资的 AI,但同时它们也会以违法的方式肆无忌惮地作弊、犯下重罪——那听上去会离谱到不行。而在当时,你可能更倾向于去谈 GPT-2 之类东西那些极其实际、直接的后果。但从某种意义上说,很多具体细节你显然是没法预见的,可事情的大致形状,你在那时候就已经可以开始推理了。只不过那样做会很难。所以我确实感觉相当困惑。但我确实觉得,重要的是——关于这档播客我一直在想的一件事是——重要的是去进行那场我当年希望自己进行过的对话。


[2:11:56] Dwarkesh Patel

The way you would have hoped you would have been talking about AIs like the present ones in 2016, rather than talking about rando bullshit about, I don't know what the topic of conversation was in 2016.I think in maybe 10 years we'll have hoped we're talking about the industrial explosion and the nature of AIs that are hard to monitor and so on.And, okay, I'll start thinking about it.Yeah.I hope that the world thinks about this in time and catches up.And I hope that the responses are good instead of bad.I don't know how optimistic I am overall, but, you know, there's good stuff to do.Yep.Cool.Thanks, Ryan.

也就是说,你会希望自己在 2016 年谈论的,是像今天这样的 AI,而不是在扯些不知所云的破事——我都不记得 2016 年大家聊的话题是什么了。我想,大概再过 10 年,我们会希望自己此刻谈的是工业爆炸、以及那些难以监控的 AI 的本质等等。那好,我这就开始琢磨这些。

Ryan: 嗯。我希望这个世界能及时想明白这件事、把认知补上,也希望应对是好的而不是糟的。我不知道自己整体上到底有多乐观,但确实有一些好事可以去做。

Dwarkesh: 对。很好。谢谢你,Ryan。