ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.168 · 全文

Why smarter AI models could drive up compute prices 10x

频道: Dwarkesh Podcast
视频: https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive-video
原文语言: en
统计: 共 11 轮 · Dwarkesh Patel 11


[0:00] Dwarkesh Patel

Today, I want to talk about what the compute situation for the labs will look like over the next few years.For the last three consecutive years, Anthropic's revenue has 10x year over year, and it's likely to do so again this year.So they ended last year with $9 billion in revenue.I think they'll probably end this year with somewhere between $100 billion to $150 billion in revenue.Now, for this trend to continue, Anthropic would need to make $1 trillion in revenue by the end of next year.Of course, there's no deep reason why this has to be true.It's a very wild conclusion, and it's ultimately a question of AI capabilities.Does AI get that useful by the end of next year?

今天我想聊聊,未来几年里各大实验室(labs,指前沿 AI 模型公司)的算力(compute)局面会是什么样子。过去连续三年,Anthropic 的收入都做到了同比 10 倍增长,今年很可能还会再来一次。他们去年年底收入是 90 亿美元。我估计他们今年年底的收入会落在 1000 亿到 1500 亿美元之间。那么,要让这个趋势继续下去,Anthropic 就得在明年年底做到 1 万亿美元的收入。当然,没有任何深层理由说这件事非成立不可。这是个非常疯狂的结论,而它归根到底是一个关于 AI 能力的问题:到明年年底,AI 真的能变得那么有用吗?


[0:30] Dwarkesh Patel

But suppose the trend does continue.Well, I want to think through what happens in that world.Now, the other big trend in AI is that lab compute only 3x year over year.For a lab to keep 10x-ing revenue year over year while compute only 3x,one of the following three things needs to happen, or some combination of the three needs to happen.One, lab margins have to increase.Two, the price of compute has to increase.Or three, the percentage of compute that lab spent on inference rather than training has to increase.My understanding is that basically all three of these things are already happening.With regards to the margins, Anthropic's inference margins reportedly went from 40% in the middle of last yearto upwards of 80% now available.With regards to compute, the spot prices for compute are more than 40% higherthan they were in the February trough that we had earlier this year.And with regards to the share of compute that goes to trading versus inference,in 2024, according to EPOC, OpenAI was spending just a quarter of its compute on inference.And that number is likely closer to 50%, if not higher now.Now, labs have deferred not to do this final thing, of increasing the share of compute they spend on inference.

但假设这个趋势确实延续下去了。我想把那个世界里会发生什么推演清楚。AI 领域另一个大趋势是:实验室手里的算力每年只增长 3 倍。一家实验室要想在算力只涨 3 倍的前提下、把收入做到每年涨 10 倍,下面三件事必须发生其中一件,或者三者的某种组合。第一,实验室的利润率(margins)必须提高。第二,算力的价格必须上涨。第三,实验室花在推理(inference,模型对外提供服务、回答用户请求所耗的算力)而不是训练(training)上的算力占比必须提高。据我了解,这三件事基本上都已经在发生了。利润率方面,据报道 Anthropic 的推理利润率从去年年中的 40%,涨到了现在的 80% 以上。算力方面,算力的现货价格(spot prices,随买随用的即时市场价)比今年早些时候 2 月那个低谷高出 40% 以上。而在训练与推理的算力分配方面,2024 年,根据 Epoch AI 的数据,OpenAI 只把四分之一的算力花在推理上;现在这个数字很可能已经接近 50%,甚至更高。不过,实验室其实一直不愿意去做最后这一件事——提高花在推理上的算力占比。


[1:38] Dwarkesh Patel

The way the labs see the world, the whole point of inference revenueis to help convince investors to give you more money in order to train the next bigger, better model.And if you're spending most of your compute on inference, you're basically declaring that AI progress has stalledand you're just now in the business of being a cloud provider.Now, this is a less compelling business than building AGI.And so the labs do not want to be in this business, nor do they think they're in this world.They think that within a year, they'll have built models that make the current ones look extremely shitty.But they need to invest a lot of their compute, the majority of their compute,into doing the training and experiments that are necessary to build the next model.So that leaves only two options for how you can get out of this gap between the fact that lab compute only increases 3x year over year,but revenue increases 10x.Either the lab's margins have to increase so that they get this surplus,or the price of compute has to increase so that everybody in the stack below the lab gets a surplus.It's not clear to me which world we end up in.Do we end up in a world where we go from 80% margins for some of the top models to greater than 90% margins

在实验室看来,推理收入的全部意义,是帮你说服投资人再多给你一笔钱,好让你去训练下一个更大更好的模型。而如果你把大部分算力都花在推理上,你基本上就是在宣告:AI 的进展已经停滞了,你现在做的不过是云服务商(cloud provider)的生意。这门生意可比「造出 AGI(通用人工智能)」没吸引力多了。所以实验室既不想做这门生意,也不认为自己身处那样的世界。他们认为,一年之内他们就能造出让现在这批模型显得烂得不行的新模型。但要做到这一点,他们就必须把大量算力——大部分算力——投进为造下一个模型所必需的训练和实验里去。所以,要弥合「实验室算力每年只涨 3 倍、收入却要涨 10 倍」这道缺口,就只剩下两个选项:要么实验室的利润率上升,让这份盈余落进实验室自己口袋;要么算力的价格上涨,让实验室下面那一整条链条上的人拿走这份盈余。我也说不准我们最终会落进哪一个世界。我们会不会走进这样一个世界:某些顶尖模型的利润率从 80% 一路涨到 90% 以上——


[2:46] Dwarkesh Patel

if the lab margin effect dominates?Well, that would require the leading model to be so far ahead of the competitionbecause the nature of margins, why they exist in a market economy,is that the thing you are serving is so much better than what somebody else could go get and replace you on the market.But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90%and they don't get competed away at that level.So that leaves only one other possibility of this escape valve between these two trends,which is that the price of compute has to increase.As I mentioned, this is already starting to happen.And the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulatebecause they can't just go out and buy a spot instance.They need to make sure that they get enough scale to get really good efficiency and flexibilityand also that they have the kind of compute that lends itself to the security they need for their own weightsand for their customers' information.So I think a relevant case study here is to look at the compute that Google and Anthropic are renting from SpaceX.Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s.

——如果实验室利润率这个效应占了上风的话?可那需要领先模型远远甩开竞争对手才行。因为利润率之所以能在市场经济里存在,本质原因就是:你提供的这个东西,比别人能在市场上另找一家替代你的东西好太多。但要我认真设想「智能」这种东西的利润率能高过 90%、而且在那个水平上还没被竞争抹平,实在太离谱了。所以就只剩下另外一种可能,来充当这两条趋势之间的泄压阀:算力的价格必须上涨。我前面提到,这件事已经开始发生了。而且,当你专门去看前沿实验室真正需要囤下来的那一批算力时,这个效应还要更强——因为他们不可能就跑出去买个现货实例(spot instance,按小时随买随用的云计算资源)了事。他们必须确保拿到足够的规模,才能获得非常好的效率和灵活性;还必须拿到那种能满足安全要求的算力,用来保护他们自己的模型权重(weights)以及客户的信息。所以我觉得这里一个值得参考的案例,是去看看 Google 和 Anthropic 从 SpaceX 那里租算力的价钱。比如 Google,每个月为 11 万块 GPU 支付 9 亿美元,这批 GPU 是 GB200 和 GB300 的混合。


[3:55] Dwarkesh Patel

The price that Google is paying here is 2x the spot price per hour for those GPUs.And that spot price itself is more than 40% higher than it would have been in February.I want to emphasize a key conclusion here, that as AI models get smarter,they'll be better able to monetize the same amount of compute.If a true human-level software engineer could run on an H100 equivalent,then at today's prices for software engineers, that H100 should rent for over $250K a year.That's over 15x the current spot price for an H100.And this is not even accounting for the fact that your AI can work nights and weekends.Of course, you might expect that if we had 10 million extra software engineers suddenly appear in the economy,the marginal value of a software engineer would decrease,and thus the revenue that that H100 would be able to generate would not be 15x higher than it is right now.But I actually don't know if this is true.If we apply this argument to people instead of AIs,then this would be the classic lump of labor fallacy.For example, economists generally believe that high school immigration does not decrease wages in the long runbecause of how innovation and specialization increase the value of labor.

Google 在这里付的价钱,是这些 GPU 每小时现货价的 2 倍。而这个现货价本身,又比 2 月份的水平高出 40% 以上。我想强调这里的一个关键结论:随着 AI 模型变得更聪明,它们把同样一份算力变现的能力也会更强。如果一个真正达到人类水平的软件工程师,可以跑在一块 H100(英伟达数据中心 GPU)等效的芯片上,那么按今天软件工程师的薪资水平,这块 H100 一年就该能租出 25 万美元以上的价钱。这是 H100 当前现货价的 15 倍还多。而且这还没算上你的 AI 可以通宵干、周末也干。当然,你可能会觉得:如果经济体里突然凭空冒出 1000 万个软件工程师,软件工程师的边际价值就会下降,于是那块 H100 能创造的收入也就不会是现在的 15 倍。但我其实不确定这是不是真的。如果把这套论证套到人身上,那它就是经典的「劳动总量谬误」(lump of labor fallacy,即误以为社会上的工作总量是固定的一块饼,多来一个人干活就少一个人有活干)。举个例子,经济学家普遍认为,(高技能)移民从长期看并不会压低工资,因为创新和专业化分工会把劳动的价值抬上去。


[5:06] Dwarkesh Patel

Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore.But if you believe what standard economics says,then the marginal value of labor, and thus the marginal value of compute,should stay astonishingly high.So let's think about what changes in such a world.One of the things that would happen is that as the top labs get better and better at monetizing compute,and the cost of compute increases, it becomes harder for anybody else to compete against thembecause they have to bid for this resource against somebody who is basically able to make better use of it.Another thing that will happen, and I think this is actually the most interesting implication of this whole thought exercise,is that if you can train the best, most efficient model,then you'll be able to charge much higher margins than you can today.This is the Alkin-Allen effect in economics,and what it's basically saying is that if it costs $20 an hour to rent an H100,then it would be extremely stupid to use a weaker, less efficient modelbecause it's going to burn more tokens on your expensive compute to get the exact same result.So labs will be able to charge a much larger premium if they can train a model

也许这一次的劳动力供给冲击会来得太大、太快,让我们没法再指望这条通用经验法则了。但如果你相信标准经济学的说法,那么劳动的边际价值——因而也是算力的边际价值——应该会保持在高得惊人的水平上。那我们来想想,在这样一个世界里有什么会改变。会发生的一件事是:随着顶尖实验室把算力变现的本事越来越强、同时算力成本又在上涨,其他任何人想跟他们竞争都会变得更难——因为你必须去跟一个「本来就能把这份资源用得更好」的人,争抢同一份资源的出价。另一件会发生的事——我认为这其实是整个思想推演里最有意思的推论——是:如果你能训练出最好、效率最高的那个模型,那么你就能收取比今天高得多的利润率。这就是经济学里的阿尔奇安–艾伦效应(Alchian–Allen effect)。它讲的基本上是:如果租一块 H100 要 20 美元一小时,那么去用一个更弱、效率更低的模型就是极其愚蠢的——因为为了得到完全相同的结果,它会在你这块昂贵的算力上烧掉更多的 token。所以,只要实验室能训练出一个更省这份稀缺投入的模型,它就能收取高得多的溢价。


[6:07] Dwarkesh Patel

that better economizes this scarce input.Basically, if you have a model that can get the same result by using less compute,then you've, in some sense, created more compute,and the value of compute is going to increase.Another thing that will happen is that a lot of current popular applications of AIwill probably get priced out.The reason AI is relatively cheap right now is thatAI just can't do a lot of things that top humans can do.But this, at some point, will no longer be the case.And at that point, Google or Anthropic or OpenAIwill be willing to pay more for the tokens to automate AI researchthan you or I will be willing to pay to make more AI slop talk.I'm a bit worried that this kind of analysis honestly pattern matches a lotonto the ways that people in the past have been wrong about scarcity.I'm, for example, thinking of the famous Simon Ehrlich bet.Paul Ehrlich was this famous doomer about population growth,and he made this bet that a basket of commodities would increase in price rather than decreasein the decade preceding 1990.And this is a very famous bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrongand how he did not anticipate the way in which market signals and human ingenuity

说白了,如果你有一个模型能用更少的算力得到同样的结果,那你在某种意义上就「创造出了更多算力」,而算力的价值也会随之上升。还会发生的另一件事是:现在很多流行的 AI 应用场景,很可能会被价格挤出局。AI 现在之所以相对便宜,是因为有很多顶尖人类能做的事,AI 还做不了。但到了某个时点,情况将不再如此。到那时,Google 或 Anthropic 或 OpenAI 为了把 AI 研究自动化而愿意为 token 付出的价钱,会超过你我为了多产出一点 AI 垃圾内容(AI slop)而愿意付的价钱。老实说,我有点担心这类分析,跟历史上人们在「稀缺」这个问题上犯错的那些套路很像。比如我就想到了著名的西蒙–埃利希赌局(Simon–Ehrlich bet)。保罗·埃利希(Paul Ehrlich)是那位著名的人口增长末日论者,他打赌说,在通往 1990 年的那十年里,一篮子大宗商品的价格会上涨而不是下跌。这个赌局非常有名,因为它被拿来说明埃利希那套马尔萨斯式的世界观错在哪里,说明他没有预料到市场信号和人类的巧思——


[7:20] Dwarkesh Patel

can find better ways to economize scarce inputs.I'm guessing that the analogy to this bet is probably wrong.Other analysis has shown that if that bet had been made in a different decade,Ehrlich might well have won.But more generally, I think the supply of compute is much less elasticand much less capable of absorbing large demand shocksand much less capable of being accommodated by using different substitutesthan the extraction of different metals.To illustrate why I think this 3x in compute capacity year over year is hard to budgeor potentially even sustain is that I don't see how any of the three elementsthat constitute that 3x can be much accelerated.So 1.4x of that is coming from Moore's Law.Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years.1.2x is coming from building new fabs.This process is ultimately going to be bottlenecked up to 2030 and potentially even beyondby just building new ASML EUV machines.Dylan, when he was on the podcast a few months ago, talked about this in great detail.And 1.8x comes from the fact that AI is absorbing a lot of wafer allocationthat was previously going to smartphones and PCs.This is probably going to hit a wall by the end of next year

——能怎样找到更好的办法去节省稀缺投入。我猜,把那个赌局类比到这里多半是不对的。另有分析显示,如果那个赌局换一个十年来打,埃利希很可能会赢。但更普遍地说,我认为算力的供给弹性要小得多,吸收巨大需求冲击的能力也差得多,能靠替代品来化解的余地同样小得多——比起开采各种金属来说是这样。要说明为什么我认为「算力产能每年 3 倍」这个数字很难往上推、甚至可能连维持都难:我看不出构成这 3 倍的三个组成部分里,还有哪一个能被大幅加速。其中 1.4 倍来自摩尔定律(Moore's Law)。别说往上加了,我觉得能把它再撑几年就已经是奇迹了。1.2 倍来自新建晶圆厂(fabs)。这个过程到 2030 年、甚至可能更久之后,最终都会卡在「造新的 ASML EUV 光刻机」这一环上。Dylan(Dylan Patel)几个月前上我节目时详细讲过这一点。剩下的 1.8 倍,来自 AI 正在吸走大量原本分给智能手机和 PC 的晶圆产能。这件事很可能到明年年底就会撞墙——


[8:35] Dwarkesh Patel

when at the leading edge N3 nodes at TSMC, AI will have gone from 60% to 86%.At some point, you have just absorbed all leading edge wafer capacity for AIand you can't keep increasing this number.So I don't know how we get even to continue to do 3x compute scaling year over yearfor the next few years, much less go beyond that.At the end of the month, I go through the time-honored tradition of closing my books.I start by opening Mercury, which is my banking platform,to make sure that all my transactions are properly categorized.Auto-categorization rules handle the predictable stuff pretty well.But I'm constantly working with new contractors, you know, tutors and researchers and videographers,and I'm also trying new tools.Manually categorizing all of these transactions would add a couple of hours of overhead every single month.So instead of going through them one by one, I have Command, which is Mercury's built-in AI.Take a stab at all of them at once.Command proposes a category for each transaction and provides its rationale.I just review, I fix anything that's off, and I approve.And once all this work is done in Mercury, it syncs everything with QuickBooks.And Command's judgment calls are genuinely good.

——到那时,在台积电(TSMC)最先进的 N3 制程节点上,AI 占的比例会从 60% 涨到 86%。到某个时点,你已经把最先进制程的晶圆产能全都吸干了,这个数字就再也涨不上去了。所以我实在不知道,接下来这几年我们要怎样才能继续做到算力每年 3 倍的扩张,更别说超过 3 倍了。

【赞助口播】每到月底,我都会走一遍那个由来已久的传统——结账关账。我先打开 Mercury,也就是我用的银行平台,确认所有交易都被正确归类了。自动归类规则处理那些可预测的项目还挺不错。但我一直在跟新的外包伙伴合作——家教、研究员、摄像师等等——我也一直在试新工具。要是把这些交易一笔一笔手动归类,每个月都会多出好几个小时的额外开销。所以我不再一笔笔过,而是让 Command(Mercury 内置的 AI)一次性把它们全都试着处理一遍。Command 会为每笔交易提出一个分类,并给出它的理由。我只要复核一下、把不对的改掉、然后批准就行。等这些活在 Mercury 里干完,它会把一切同步到 QuickBooks。而且 Command 的判断确实相当到位。


[9:38] Dwarkesh Patel

It does the obvious things like looking at the vendor,but it also investigates who on my team made the purchaseand looks at notes and memos to build up as much context as possible.This is just one of the ways you can use Command to automate the back end of your business.To learn more, go to mercury.com slash command.Mercury is a fintech company, not an FDIC-insured bank.Banking services provided through Choice Financial Group and Column N.A. members FDIC.AI-generated responses and suggested actions may vary and are not guaranteed.Now, I want to clarify that at some point in the future, compute will get cheap again.At some point, we'll just have robots that can convert shores of silica sandand mines of copper into new computer chips.And then the price of compute is basically the raw inputsand the tools required to do this processing.I'm just talking about this current pre-singularity regimewhere AI compute merely 3x's year over year,which is not enough to offset how much more valuable AI is becoming over time.By the way, the fact that anthropic revenue has been 10x in year over year,whereas their compute has only been 3x in year over year,I think illustrates how strong the economies of scale are in the model business.

【赞助口播(续)】它会做那些显而易见的事,比如看供应商是谁;但它还会去查是我团队里的谁做了这笔采购,会去看备注和留言,尽可能把上下文补齐。这只是你能用 Command 把公司后台自动化的方式之一。想了解更多,请访问 mercury.com/command。Mercury 是一家金融科技公司,不是 FDIC 承保的银行。银行服务由 Choice Financial Group 与 Column N.A. 提供,二者均为 FDIC 成员。AI 生成的回复和建议操作可能有出入,不作保证。

好,我想澄清一点:在未来的某个时点,算力会重新变便宜。到某个时候,我们就会有机器人,能把成片的硅砂海岸和铜矿变成新的计算机芯片。那时候算力的价格,基本上就等于原材料,加上完成这套加工所需的工具。我这里讲的只是当下这个「奇点前(pre-singularity)」的阶段——在这个阶段里,AI 算力仅仅是每年 3 倍,而这不足以抵消 AI 本身正变得越来越值钱。顺便说一句,Anthropic 的收入每年 10 倍、而它的算力每年只有 3 倍,我觉得这件事本身就说明了模型这门生意里的规模经济(economies of scale)有多强。


[10:43] Dwarkesh Patel

And logically, this makes sense.When you train a model, you just have to spend this one-time costto learn all these different skills that then get to be shared across all your users.This is very unlike human labor, where each instance has to be retrained from scratch.I wish we didn't live in a world with such strong economies of scale for intelligencebecause I'm worried about power concentration.But it seems we do.Okay, this was a narration of a blog post that I also released on my website at dhwarkesh.com.Check it out for other posts or to be notified when I release a post in the future.Otherwise, I'll see you for the next full episode.Thank you.

从逻辑上讲这也说得通。当你训练一个模型的时候,你只需要付一次这笔一次性成本,去学会所有这些不同的技能,之后这些技能就可以被你所有的用户共享。这跟人类劳动非常不一样——人类的每一个个体都得从零开始重新培训一遍。我倒宁愿我们不是活在一个「智能的规模经济如此之强」的世界里,因为我担心权力集中的问题。但看起来我们就是活在这样的世界里。好了,这一期是我在自己网站 dwarkesh.com 上发布的一篇博文的朗读版。去看看我的其他文章吧,或者订阅一下,以后我发新文章时你会收到通知。除此之外,我们下一期完整节目再见。谢谢。