Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden
频道: Sequoia Capital
视频: https://www.youtube.com/watch?v=vPnVTHYplrQ
原文语言: en
统计: 共 66 轮 · Host 18 · Angela 28 · Katelyn 10
[0:00]
The last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution, you have coordination. And at the coordination layer, we're beginning to think of these things called like strategies where basically it's almost like a meta harness. The true low-level harness is designed for execution. But the next one is about okay if tokens aren't really fungeible and you need to give them different jobs like maybe some this token is advising versus this token is executing you want to start composing these like these kind of orchestrated strategies that go together and they should sit on top of all these things because at the end of the day you still need to execute and the execution still needs to know what to do. So everything in theory should kind of like ladder together. And so I think you know if you were to look at our road map and the maybe kind of project forward a little bit where you kind of expect us to go. We'll move more and more from the knowledge layer to the execution layer and from the execution layer to the kind of coordination layer in terms of the abstractions that [music] you can see us put out. [music] Caitlyn and Angela, thank you so much for joining us today. Lauren and I are thrilled to have you here. You are responsible for building Anthropics platform and so you are responsible for building what I think is one of the most important if not the most important developer platform in the world and we are really excited to interview you today to understand more about what's ahead and so maybe just to get started can you give us the context of you know what is anthropic platform and where do you sit within anthropic? Yeah. So platform is both our externally facing APIs, our developer platform that people build on top of when they want to build applications and system that systems that access cloud's intelligence as well as internally um we run our product infrastructure and basically we're the layer that our apps build on top of uh internally as well.
在这之上、最后一层抽象,大概就是协调层(coordination layer)了。所以你有知识层、有执行层、还有协调层。在协调层这一层,我们开始琢磨一种叫 strategies(策略层)的东西——它基本上就像一个 meta harness(元 harness)。真正的底层 harness 是为执行而设计的,而上面这一层要解决的是:既然 token 并不是完全可以互换的、你得给它们分派不同的活儿——比如这个 token 负责出谋划策、那个 token 负责实际执行——那你就会想把这些编排好的 strategies 组合到一起。它们应该叠在所有这些能力之上,因为归根结底你还是得执行,而执行环节仍然需要知道该做什么。所以理论上,所有这些东西应该像梯子一样层层咬合、递进上来。所以我想,如果你去看我们的路线图、稍微往前推演一下我们会往哪儿走:在对外释放的抽象层面,我们会越来越多地从知识层走向执行层,再从执行层走向协调层。
【主持人】Katelyn、Angela,非常感谢两位今天来做客。我和 Lauren 都特别高兴能请到你们。你们负责搭建 Anthropic 的平台——也就是说,你们负责打造的这个东西,在我看来是全世界最重要的开发者平台之一,甚至可能就是最重要的那一个。今天我们特别想采访你们,多了解一下接下来会有什么。那不如就从头开始:能不能先给我们讲讲背景——Anthropic 的平台到底是什么,以及你们在 Anthropic 内部处于什么位置?
【嘉宾】好的。所谓平台,一方面是我们对外的那套 API,也就是开发者平台——大家想构建应用、想构建能调用 Claude 智能的系统时,就在它上面搭建;另一方面,在内部,我们也负责运营我们的产品基础设施。基本上,我们就是 Anthropic 自家应用在内部所依托的那一层。
[1:49] Host
Awesome. What's your northstar as a team?
太棒了。作为一个团队,你们的 north star(北极星指标)是什么?
[1:51] Angela
It's a great question. We actually because we have both internal and external. We actually kind of have like two north stars which is probably like you know you' be like why there should only be one north star but um no we
这个问题问得好。其实呢,因为我们同时有对内和对外两块业务,我们算是有两个 north star。你可能会说——north star 不是应该只有一个吗?但对我们来说不是,我们……
[2:00]
different planetary system.
属于不同的行星系嘛。
[2:02] Angela
Yes exactly they're separate solar system so it's fine. Um but uh on the internal side like we really want to provide is like literally as much leverage as possible for our internal teams to be able to ship like AGI pill products. Um and we want them to be able to move fast, be able to have reliable like great like uh platform uh to be able to build on top of. But I think that key bit about speed is like really intentional for us and we really really care about that internally. Externally um we actually have a lot more like complicated set of things. Um but one of the true norths that we have there is to be able to basically give any builder the tools to be able to work with claude to build whatever they want to build. And so it's a bit of a broad statement but as a result that uh boils it itself down into you know being wherever that business is. Like we really care about like bringing our platform really really close to that business. This is why we spend a lot of time with the hyperscalers integrating really closely uh directly with them like AWS, Google, so on and so forth. Um and it is a lot of like primitives that we end up creating. We want people to be able to express what they think their product should be. We want them to be able to almost do like custom software in their own way. You know, like in this new world with AI, uh what used to be probably economically impossible was that last mile of of custom software now in theory should be like very very achievable. Um, and we want to give them all the tools and all the capabilities to go and do that. And so sometimes that comes in the form of primitives and APIs and higher order abstractions. And sometimes that comes in the form of just like standards. Uh, so for example like skills and MCP. Um, those are things just like cloud needs them to be useful and we can just give them out to the rest of the ecosystem, work with everyone to help you create those things and get the best out of cloud.
对,没错,它们是两个各自独立的太阳系,所以没关系。不过呢,在对内这一侧,我们真正想提供的,说白了就是尽可能给内部团队最大的杠杆,让他们能把真正押注 AGI(AGI-pilled)的产品做出来、发出去。我们希望他们能跑得快,希望他们有一个可靠、好用的平台可以在上面搭建。而“快”这一点对我们来说是非常刻意强调的,我们内部真的非常非常看重速度。对外这一侧,我们要面对的东西其实要复杂得多。但我们对外的一个真正的“北极星”,是要让任何一个 builder 都能拿到工具、能用 Claude 去构建他们想构建的任何东西。这么说是有点宽泛,但正因如此,它落下来就变成了一句话:业务在哪儿,我们就去哪儿。我们真的很在意把我们的平台带到离那个业务非常非常近的地方。这也是为什么我们花了大量时间跟那些超大规模云厂商(hyperscalers)做非常紧密、非常直接的集成,比如 AWS、Google 等等。最终我们会创造出很多 primitives(基础原语)。我们希望大家能够表达出他们心目中产品应该是什么样子,希望他们能以自己的方式去做类似“定制软件”的东西。你知道,在这个有了 AI 的新世界里,过去在经济上大概根本不可能做成的那“定制软件的最后一公里”,现在理论上应该变得非常非常可行了。我们想把所有的工具、所有的能力都交到他们手上,让他们去做这件事。所以有时候,这体现为 primitives、API 和更高阶的抽象;有时候,它就只是以“标准”的形式出现。比如说 skills 和 MCP——这些东西本来就是 Claude 自己需要、才有用的,那我们干脆把它们开放给整个生态,跟大家一起来帮你把这些东西做出来、把 Claude 的能力发挥到极致。
[3:52] Angela
So um I would say externally you know we really are oriented around just helping you just be able to build but internally that orientation while still existing is probably more you know specified towards speed and being able to move really quickly.
所以总的来说,对外,我们真的是围绕着“帮你把东西做出来”来定位的;而对内,这个定位虽然同样存在,但更多是聚焦在速度上——聚焦在能不能跑得非常快。
[3:54] Host
How do you decide what goes into the platform what gets externalized and what doesn't to decide what products should be available?
那你们是怎么决定什么东西进平台、什么东西对外开放、什么不开放的?也就是怎么决定哪些产品该对外提供?
[4:01] Angela
Yeah I mean we generally try to have a philosophy that we try to be consistent across the board. It's actually one of the reasons uh why we do internal and external. Um there's plenty of other you know platform businesses and constructs where you actually like bifrocate these two things. Um for us we kind of try to intentionally keep it equal and then as a result we try to hold this philosophy as much as we can around like you know for any builder internal or external even though if our internal builders might have some slightly different requirements in the same way any user would have slightly different requirements. Uh we want to have the same primitives that are available to everyone. And one of the maybe the overarching thesis for that is that we've just seen like the capabilities of these models just grow and such just exponential and it's really hard to figure out like a longlasting form factor. I think two years ago we were all like everything's chat and now everyone's like forget chat and just like agents and like there's going to be another form factor, another form factor. Um and we kind of imagine that like constantly evolving. And so the best way for us to kind of enable that for everyone and also ourselves is to actually build a really robust platform that gives people those kinds of like tools to figure out what those form factors are. And I don't think we by any means feel like we're the only ones capable of figuring out that form factor like not at all. In fact, the more democratization we can do on that and help people and allow people to experiment, I think the the more those form factors will actually kind of naturally come out of the market.
嗯,总体上我们努力秉持一种理念:尽量在方方面面保持一致。这其实也正是我们为什么把对内和对外放在一起做的原因之一。业界有很多别的平台型业务和架构,是会把这两件事拆开、分而治之的。而我们是有意地尽量让两者对等,然后据此尽可能守住这样一个理念——面对任何一个 builder,无论对内还是对外,即便我们内部的 builder 可能会有些略微不同的需求(就像任何一个用户都会有些略微不同的需求一样),我们也希望提供给所有人的是同一套 primitives。这背后一个也许是统领性的判断是:我们眼看着这些模型的能力就这么指数级地往上涨,你真的很难判定出一种能长久稳定的产品形态(form factor)。两年前我们都还觉得“一切皆是 chat”,现在大家又都说“别提 chat 了,全是 agent”,而且往后还会有下一种形态、再下一种形态。我们心里预设的就是它会一直这么演化下去。所以,无论是给别人、还是给我们自己,我们能做的最好的事,就是实实在在搭一个非常稳健的平台,把这类工具交给大家,让他们去摸索出那些新形态到底是什么。我们绝不觉得只有我们才有本事想明白那个形态,完全不是。事实上,我们在这件事上能做的民主化越多、能帮到并放手让越多的人去实验,这些形态其实反而越会自然而然地从市场里冒出来。
[5:15] Katelyn
Yeah. Yeah. And I think within our team, we've we've had moments where we're experimenting even with just like a packaging up of our primitives in a different sort of higher order way. And we've thought about, okay, cool. We've solved this exact type of problem with this product that we've built into the world. And so we can go and dog food it for ourselves, but we'd never want to fall into this trap of like we're overindexed on the problem as it needs to be solved for an internal user. Like because exactly what Angela said, internal users have very specific requirements. External users have very specific requirements and so if you overindex on one or the other, you fall into a trap. So a lot of the time what we'll do is dog food something internally at the same time that we open up early access of some sort with external customers so that we can kind of get a range of feedback and bring those things back into the platform.
对,对。而且我觉得在我们团队内部,我们也有过这样的时刻:我们甚至会去尝试把 primitives 用另一种更高阶的方式打包起来。我们会想——“好,酷,我们做出来、推向世界的这个产品,恰好解决了这一类具体问题,那我们就可以自己拿来 dogfood(内部试用)。”但我们绝不想掉进这样一个陷阱:过度围着“某个内部用户需要被解决的问题”来打转。因为正如 Angela 说的,内部用户有非常具体的需求,外部用户也有非常具体的需求,你要是过度偏向其中任何一边,就会掉坑里。所以很多时候我们的做法是:一边在内部 dogfood 某个东西,与此同时也向外部客户开放某种形式的早期访问(early access),这样我们就能拿到一个范围更广的反馈,再把这些东西带回到平台里去。
[6:00] Host
I'd love to talk about the higher levels of abstraction that you discussed. So I guess at the base level this is just you know raw access to CLA opus or whatever tokens. How do you think about the I guess the layer cake of abstractions above that?
我很想聊聊你们提到的那些更高层的抽象。我理解,在最底层,这无非就是对 Claude Opus 之类模型的原始(raw)访问、拿到 token。那你们是怎么看待在这之上那一层层叠起来的“抽象千层蛋糕”的呢?
[6:14] Katelyn
Yeah, if you look back um so when I joined Anthropic around a year ago um the platform was basically just the messages API. It was a messages API. Um you know we had come out with standards like MCP. We obviously have developer tooling around our SDKs and our docs and our console and things like this. But for the most part it was a stateless API. Um, and what's interesting to Angela's point on form factors evolving over time is we found a lot of our customers solving the same problems over and over again that we also were solving over and over again around as the models got better at running for longer and working with more contexts at a given time. You want to build agents that can succeed in a kind of longunning context and even a remote context that doesn't necessarily have a human in the loop. And so we found that we could piece together our primitives and stand up all the same infrastructure that we're finding ourselves standing up internally to power our own products and arrive at some higher order abstractions that let you do more agentic work out of the box. And the problems that we're solving for you are, you know, infrastructure being kind of a hard thing to deal with. like how do you figure out spawning sandboxes that are going to have the right governance and security and like you know spin them up and spin them down when you need to or the storage around transcript sessions so that you can resume a session if you stop it and pick it back up later. Um so that infrastructure is a big thing that we wanted to be able to provide more of out of the box and we do more of that today. Um, and then the second thing just being harnesses and harness engineering. There's a lot of thought and energy going into how do I do my prompt caching and how do I manage my context window as well as how do I actually just get more intelligence out of the model um, and how do I manage my costs and things like that.
嗯,回头看一下——我大概一年前加入 Anthropic 的时候,平台基本上就只是一个 Messages API,就一个 Messages API。当然,我们已经推出过像 MCP 这样的标准,也有围绕 SDK、文档、console 之类的开发者工具。但绝大部分情况下,它就是一个无状态(stateless)的 API。有意思的是——呼应 Angela 讲的“形态会随时间演化”——我们发现很多客户在反复解决同样的问题,而这些问题也正是我们自己在反复解决的:随着模型越来越擅长长时间运行、越来越能在同一时间处理更多的上下文,你就会想构建那种能在长跑式(long-running)场景、甚至在没有人类实时介入的远程(remote)场景里跑成功的 agent。于是我们发现,我们可以把手里的 primitives 拼装起来,把我们内部为驱动自家产品而反复搭建的那整套基础设施立起来,从而得到一些更高阶的抽象,让你开箱即用就能做更多 agentic 的工作。而我们替你解决的那些问题,说白了就是——基础设施本身是件很难对付的事。比如你怎么搞定 sandbox 的按需拉起:既要有正确的治理和安全策略,又要能在你需要的时候把它拉起来、用完再关掉;再比如围绕会话记录(transcript session)的存储,好让你在中断一个会话之后,之后还能恢复、接着往下跑。所以这类基础设施是我们很想更多地开箱即用提供给大家的一块,如今我们在这方面做得也更多了。第二件事就是 harness 以及 harness engineering。这里面有大量的思考和精力:我该怎么做我的 prompt caching、怎么管理我的上下文窗口,以及我到底怎么从模型里榨出更多的智能、怎么控制我的成本,诸如此类。
[8:08] Katelyn
So we've kind of packaged up our primitives a bit more in tune with the problems that we found ourselves solving to provide more of these things out of the box for people so that they can if they're building systems for themselves internally, if they're building products, they can just be more focused on the problems that they want to be solving and if they want to offload some aspects of those problems to us, they can. Um, and that's kind of the ethos. And are your customers generally choosing to opt from the grab bag of stuff that you offer or are they like how how often are they opting into the just the the managed agents offering? I guess just take care of it all for me.
所以我们把 primitives 做了一些打包,让它更贴合我们发现自己在反复解决的那些问题,从而把更多这样的能力开箱即用地提供给大家——这样,无论他们是在给自己内部搭系统,还是在做产品,都能更专注在他们真正想解决的问题上;而如果他们想把其中某些环节的问题卸给我们来扛,也可以。这大致就是我们的理念。
【主持人】那你们的客户一般是怎么选的——是从你们提供的这一堆东西里自己挑(grab bag,自选菜单式),还是说他们会有多大比例直接选那个托管式 agent 方案,也就是“这些你全帮我搞定就行”?
[8:29] Angela
Um it varies by like the the user group. So like for I would say you know like really AI native startups like the ones who are like tinkering and like experimenting at a really low layer they're just going to go for the primitives. Um and then for everyone else and these are kind of classic like more like enterprises or areas where it's like the purpose of the startup or the philosophy behind the startup isn't necessarily to optimize on um some kind of hill climbing piece. It's more like stringing together a bunch of workflows and you know providing unique user value at uh to that user. For those people um you know it's just kind of not their core competency. It's not where they want to focus their time and resources and they reach much more for these kind of like higher order like package offerings.
嗯,这要看用户群体。比如说,那种真正 AI 原生的初创公司——就是那些喜欢在很底层去捣鼓、去实验的——他们会直接冲着 primitives 去。而其他所有人呢——这些更像是典型的、偏企业级的客户,或者说这家初创公司的目的、它背后的理念本身并不是要在某种“爬坡优化(hill climbing)”上死磕,而更多是把一堆 workflow 串起来、给它的用户提供独特的价值——对这些人来说,这(底层优化)就不是他们的核心竞争力,不是他们想投入时间和资源的地方,于是他们会更多地去够那些更高阶的、打包好的方案。
[9:07] Host
What are some examples of the primitives you've released at different layers in the last few months? We've seen a few of them. Would love to hear.
过去这几个月,你们在不同层级上发布过哪些 primitives,能举几个例子吗?我们也见过其中一些,很想听你们讲讲。
[9:13] Angela
Yeah. I think maybe one framing I would give um for some of the constructs that Caitlin was talking about is like and this is a bit of an oversimplification, but effectively there's approximately like three like layers of this cake. At the very bottom is just kind of like like knowledge. And so at this layer like in many ways it's it's knowledge about the model. It's knowledge about the things that the model needs. And it's just like the ability to know how to actually do something with claude is maybe the way I'd phrase that. And so there the primitives that we have spent more and more time on uh have been actually things of the past because like we still evolve them but they tend to be a little bit more baked. Like for example there's very specific shapes and parameters we put on the messages API and it's more like trying to expressly like uh showcase Claude's like design like Claude the model's uh actual design the way it thinks the way it respects certain parameters the way it kind of like um will do tool calls like all of those different pieces. And then we started standardizing like tools. And then we started standardizing bits and pieces of like context that you could put in at different moments in time which is concretely like skills and like memory. And so those are like the kind of like knowledge layer type of abstractions that we've put out over the past um I guess like year plus plus a bit. Um the next layer of abstraction that we've actually started to spend more and more of our time on is like once you kind of know stuff you then need to like execute. And so at the execution layer, that level of abstraction is the part that Caitlin was talking about around like we're doing these like higher order pieces, but like what are we putting higher order there? It really is because you're now getting Claude to execute work. It's not just to know something, right? I can give it a question and give me an answer. You can put string a lot of that stuff together. Uh but now if you need to execute like do work, give me the output, edit files in a bunch of different systems, that becomes a lot more complicated and requires infrastructure to handle.
好。对于 Katelyn 刚才讲的那些构件,我想给一个或许有用的框架——虽然这么说有点过度简化:基本上这块“蛋糕”大概可以分成三层。最底下这一层,就是“知识(knowledge)”。在这一层,很多时候它其实是关于模型本身的知识、是关于模型所需要的那些东西的知识,或者用我的话说,就是“知道到底该怎么用 Claude 把一件事做出来”的能力。在这一层,我们越来越多投入精力的那些 primitives,其实反而是些比较“老”的东西——因为我们仍在迭代它们,但它们往往已经比较定型、比较成熟了。举个例子,我们在 Messages API 上放了非常具体的结构和参数,它更像是要明确地把 Claude 的设计呈现出来——是 Claude 这个模型本身的设计、它的思考方式、它对某些参数的遵循方式、它做 tool call 的方式,所有这些不同的部分。接着我们开始把 tools 标准化,再接着我们开始把那些可以在不同时刻塞进去的一块块“上下文”标准化——具体来说就是 skills 和 memory。所以这些就是过去大概一年多里我们推出的、属于“知识层”的那类抽象。而下一层抽象——我们其实已经开始把越来越多的精力投进去了——是:当你大致知道该怎么做之后,你接着就得去“执行(execute)”。在执行层,那一层抽象正是 Katelyn 刚才讲的、我们在做的那些更高阶的东西。但我们把什么放到了“更高阶”的位置上?归根结底,是因为你现在是在让 Claude 去执行工作了,而不只是让它“知道”某件事,对吧?我可以给它一个问题、让它给我一个答案,你也可以把很多这样的东西串起来。但一旦你需要它去执行、去干活——给我产出、在一大堆不同系统里改文件——事情就复杂多了,而且需要基础设施来支撑。
[10:56] Angela
And so that layer is basically I would say a low-level harness plus manage infrastructure as like the set of abstractions. Today we just like our highle product for that is called cloud manage agents. Um and so that's like a piece but we started to wrap more and more pieces in that. Um I think there's going to be a layer like on top of that we have like some inklings of it we started to build towards but the last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution you have coordination. And at the coordination layer, um, we've started to expose some of these in ways that like aren't very obvious, but we're beginning to think of these things called like strategies where basically it's almost like a meta harness, right? The harness, the true low-level harness is designed for execution. But the next one is about okay if tokens aren't really funible and you need to give them different jobs like maybe some this token is advising versus this token is executing this token is dreaming versus this token's executing so on and so forth you want to start composing these like these kind of orchestrated strategies that go together and they should sit on top of all these things because at the end of the day you still need to execute and the execution still needs to know what to do so everything in theory should kind of like ladder together and so I think you know if you were to look at our road map and the maybe kind of project forward a little bit where you kind of expect us to go we'll move more and more from the knowledge layer to the execution layer and from the execution layer to the kind of coordination layer in terms of the abstractions that you can see us put out.
所以那一层,基本上我会说,就是“一个底层 harness + 托管基础设施”这样一套抽象。今天我们在这方面最高阶的产品,叫 Claude Managed Agents。它是其中一块,而我们开始把越来越多的东西往里裹。我觉得在它之上还会有一层——我们对它已经有了一些朦胧的想法、也开始朝那个方向搭了一点——但这之上最后一层抽象,大概就是协调层(coordination layer)。所以你有知识、有执行、有协调。在协调层,我们已经开始以一些不那么显眼的方式把其中一部分暴露出来了,我们开始在琢磨一种叫 strategies(策略层)的东西——它基本上就像一个 meta harness,对吧?那个真正的底层 harness 是为执行而设计的,而上面这一层要解决的是:既然 token 并不是完全可以互换的、你得给它们分派不同的活儿——比如这个 token 负责出谋划策、那个 token 负责执行,这个 token 负责“做梦(dreaming)”、那个 token 负责执行,等等——那你就会想开始把这些编排好的 strategies 组合起来。它们应该叠在所有这些能力之上,因为归根到底你还是得执行,而执行环节仍然需要知道该做什么。所以理论上,所有这些东西应该像梯子一样层层咬合、递进上来。所以我想,如果你去看我们的路线图、稍微往前推演一下我们会往哪儿走:在你能看到我们对外释放的抽象层面上,我们会越来越多地从知识层走向执行层,再从执行层走向协调层。
[12:12] Host
That's a really cool.
这真的太酷了。
[12:13] Host
How do you think this all comes together into a broader ecosystem beyond just the things that you guys are building? How do you help support people building products on top of it and how do you help them get the most out of all these pieces?
那你们觉得,这一切最终会怎么汇聚成一个更大的生态——一个超出你们自己所构建的那些东西的生态?你们会怎么去支持那些在它之上做产品的人,又会怎么帮他们把所有这些部件的价值发挥到极致?
[12:25] Angela
Yeah, I think this is like super top of mind for us. like we really want to find a way to be to support as many people in doing this as we can. I think we're still like learning like a lot of the industry like has evolved. We've seen, you know, a lot of different pieces um get spun up and spun down. And I think the the operative part for for Caitlyn and I has been in the category of like making sure at least at the base layer that we provide as many primitives across the board as possible. So you know this kind of like yeah like knowledge execution coordination layer we want to give all of that out to everyone so that people can start to compose and create on top of that. Um and that's just from like I think pure builder kind of point of view. Then there's a point of view around like how do you kind of like plug in with us right like we're also building firstparty products of our own. We've also created some ways to embed natively with us like for example connectors which are built on top of the MCP spec. Um, and we try to be more open about those types of things. And we're starting to figure out like what are the right bits and pieces, but what we're really trying to do is get to a place where, you know, in a company is able to get created and built on uh they can build whatever products that they want. They can build agents if they need to. And then those agents and those products could be things that could plug into other agents. Some of those agents could be cloud agents, some of those agents could be other people's agents. Um, but we want to be able to enable that kind of like transactability across the board. And then I think in order for all of that to kind of ultimately be true, there is a bit around like standard setting and I think there's the traditional standard setting which is around you know how do systems interoperate um and that's uh you know things that you've kind of seen us do with like skills and MCP but they're at again like the builder layer.
是啊,这件事其实一直是我们心里最上心的。我们特别想找到办法,尽可能支持更多人去做这件事。我觉得我们也还在摸索,整个行业变化太快了——你看,各种东西不断地被造出来,又不断地被淘汰。对 Katelyn 和我来说,我们真正在发力的方向,是至少在最底层,尽量把各个环节的 primitives 都提供出来。就是说,像知识、执行、协调这几层能力,我们想把它们全都开放给所有人,让大家能在这之上去组合、去创造。这纯粹是从一个 builder 的视角出发的。然后还有另一个视角,就是你怎么跟我们对接——毕竟我们自己也在做 first-party 的产品。我们也做了一些能跟我们原生嵌合的方式,比如 connectors,它就是建在 MCP spec 之上的。这类东西我们尽量做得更开放。我们也还在琢磨到底哪些环节该怎么切,但我们真正想达到的状态是:一家公司成立后,他们想 build 什么产品就 build 什么产品,需要的话可以做自己的 agent,而这些 agent、这些产品又能接到别的 agent 上去——有的可能是 Claude 的 agent,有的可能是别人家的 agent。我们希望让这种跨系统的互通、可交易的能力在全局都成立。而要让这一切最终真正成立,还得有一块是关于标准制定的。传统意义上的标准制定,是关于系统之间怎么互操作——就是你看到我们在 skills、MCP 上做的那些事,不过那些还是在 builder 这一层。
[14:08] Angela
I think at a higher order layer there's also a bit around interoperability and standard setting around how do we all kind of like treat safety together and you know we've talked to a lot of these companies and this is less from you know philosophies aside just more like no one really wants to have technology that's like for example like doing negative things on um on their service right so cyber I think is a great example of this uh you want to protect your own systems from like negative actors or or bad actors and so like these kinds of like standard settings of like how we find ways to partner with more and more people to be like, yeah, we all kind of want to make sure our critical infrastructure is good. We all want to prevent like fraud or any of those things from happening and how can we work better with each of these members. I think on the last layer, we're still kind of like we're still evolving and I think we're still very much like trying to find ways that we can be better and work with the rest of the industry to bring people along and and work with them. Um but those are kind of like you know the higher order primitives or pieces that we wish to kind of like be in place. Um so they can work with folks to to ultimately solve this. I think if I were to like take a step back at the end of the day on on all of these things um you know like this technology is so transformative. Uh and if it's a little bit like electricity in the sense like before electricity there was just like you know you had to like have a candle and it was like you can only do so many things. Um, but with electricity, the reason why it's such a transforming technology for all of us and so greatly of a utility is because you can actually like wire it into everything. Everyone is able to actually access it. We also have like standards and ways to plug in and do all the pieces that we need. And that's not something that anybody can do by themselves. They always have to work with the ecosystem and work with partners um to figure out a path forward.
我觉得在更高一层,还有一块是关于互操作和标准制定的——就是我们大家怎么一起对待安全这件事。我们跟很多这样的公司聊过,这其实不太是什么理念之争,更实际的是:没有人真的希望有一项技术,比如说在自己的服务上干一些负面的事,对吧?网络安全就是个很好的例子——你肯定想保护自己的系统不被那些恶意的、搞破坏的人利用。所以这类标准的建立,就是我们怎么找到办法跟越来越多的人合作,大家一致地说:对,我们都想确保关键基础设施是稳的,都想防止欺诈之类的事发生,那我们怎么跟这里的每一方更好地协作。至于最后这一层,我们其实还在演进,也还在很努力地找办法做得更好、跟整个行业一起把大家带上路、跟他们一起干。不过这些差不多就是我们希望能到位的那些更高阶的 primitives、或者说环节,好让大家能一起最终把这个问题解决掉。如果让我在所有这些事情上退一步看的话——这项技术真的是变革性的。它有点像电:在有电之前,你就只能点根蜡烛,能干的事情非常有限。而电之所以对我们所有人来说是这么一项变革性的技术、这么大的一种基础设施红利,是因为你真的能把它接进一切东西里,人人都能用上它。我们还有标准、有各种接入的方式,把所有需要的环节都做好。而这不是任何一家能自己独立完成的,你总得跟整个生态合作、跟合作伙伴一起,才能摸索出一条路来。
[15:39] Host
How do you think about the philosophy of building an open ecosystem uh versus a walled garden? And you know, how do you think about what products are really important for you to own first party versus where you're perfectly happy to plug into other components of the ecosystem?
你们是怎么看待「建一个开放生态」和「建一个围墙花园」这两种理念之间的取舍的?还有,你们怎么判断哪些产品是你们非得自己做 first-party 不可,哪些则是你们很乐意直接接入生态里别人已有的组件?
[15:54] Katelyn
Yeah, there's so maybe in using Angela's kind of layered cake that we talked about a little bit earlier, you'll see that on some pieces of this like execution for example, um what we've done within something like cloud managed agents and I think over time you'll see us try to make this a little bit more modular. We actually aren't precious about you should run these things on our infrastructure like it should be sandboxes that we control or it should be a storage layer that we control. Um well we actually like for example we launched self-hosted sandboxes and we partnered with modal and versel and cloudflare and a bunch of other folks um even like Amazon's new microVMs um to have a first class offering where you can go plug any of those things in. Um, we launched MCP tunnels so that you can call out to your MCP servers that are behind your firewall, right? And um, be able to punch through there. And so for some of these things, we, you know, the weather, whether it runs on our infrastructure versus somebody else's infrastructure is actually not important to us because the thing that's important to us is more that the architecture of how you put together these agents in a way that will be powerful, in a way that will be reliable and scalable. um we have strong opinions on that and you can kind of just conform to the interfaces that we put out there and plug those things in. Um and we think that that generally is a thing that works really well. Yeah, I think on the the kind of like verticals where we might build products um you know I think we we kind of have like two frames here. The first one is we are always trying to figure out a form factor like an evolving form factor. We by the way don't think form factors are like static. It's like a dynamic thing. So what might be awesome for one year's worth of AI development will probably not be awesome for the next year's worth. And we just kind of try to have that mentality. We tell the team uh just overall like around anthropic. Everyone's always trying to be like is this agi pill enough?
嗯,那我借用一下 Angela 前面讲的那个「分层蛋糕」的说法吧。你会发现,在其中一些环节上,比如执行这块,我们在 Claude managed agents 里做的那些东西——而且随着时间推移,你会看到我们尽量把它做得更模块化。我们其实并不执着于「这些东西必须跑在我们的基础设施上」,比如非得是我们控制的 sandbox、非得是我们控制的存储层。相反,举个例子,我们上线了 self-hosted sandboxes,还跟 Modal、Vercel、Cloudflare 等一大批伙伴合作,甚至包括 Amazon 新出的 microVMs,做成一个一等公民级别的能力,你可以随便把这些东西接进来。我们还上线了 MCP tunnels,这样你就能调用你防火墙后面的 MCP server,把那层打通。所以对这些东西来说,它到底跑在我们的基础设施上、还是跑在别人的基础设施上,对我们其实并不重要——因为对我们更重要的是:你把这些 agent 组装起来的架构,得是强大的、可靠的、可扩展的。在这一点上我们有很强的观点,你基本只要照着我们放出来的那些接口去对接、把东西插进来就行。我们觉得这套东西整体上是很好用的。至于那些我们可能会自己做产品的垂直领域,我觉得我们大概有两个框架。第一个是,我们一直在琢磨一种 form factor——一种会不断进化的 form factor。顺便说一句,我们不认为 form factor 是静态的,它是个动态的东西。所以某一年的 AI 发展里很棒的形态,到了下一年多半就不再那么棒了。我们就是抱着这种心态。我们跟团队说——其实是整个 Anthropic——大家总是在互相问:这个东西够不够「AGI-pilled」(够不够 AGI 信仰)?
[17:47] Angela
Um and then we always have this mentality of like you know we built something it works it was cool for a year and maybe it's not the right next thing and so throw it away try again. Um and we we tell like platform users the same thing. um I just think that's probably just like you know attached to the technology but so yeah one one principle is like trying to always constantly find this new form factor. So sometimes we'll like launch products in certain areas to try to showcase a new type of form factor. Um, it's not necessarily because we think it's like the biggest ham or the most important thing to go after, but sometimes you're like, okay, this is like always been a really difficult thing and people have always communicated this way or tried some things this way and can we show that maybe there's a slightly different way. Um, and because the model capabilities are are so advanced now, can we try to express it a bit differently? Um,
然后我们一直有这么一种心态:我们做出来个东西,它能用、火了一年,也许它就不再是下一步该做的对的东西了——那就把它扔掉、重来。我们也跟平台用户这么说。我觉得这大概就是这项技术本身的特性决定的。所以,一条原则就是不停地去找这种新的 form factor。有时候我们会在某些领域上线产品,就是为了展示一种新型的 form factor。这倒不一定是因为我们觉得它是块最大的肥肉、或者是最该去攻的事,而是有时候你会想:好,这件事一直特别难搞,大家一直都是用某种方式沟通、或者一直用某种方式在尝试——那我们能不能证明,也许有一种稍微不一样的路子?而且既然现在模型的能力这么强了,我们能不能把它换一种方式表达出来?
[18:24] Host
what's an example of that?
能举个例子吗?
[18:25] Angela
Yeah, you know, like uh cloud design is a little bit of of that way. I think depending on how you squint, you might see it as like a way that we kind of going into design as as like you know one of the verticals. But more often than not, it's like if you take a look at what we're trying to do with that product, there's a couple of like decisions that were made in there. The first one is that like you can actually try to offload more and more and more to Claude. Um, and so it tries to be kind of opinionated on like, you know, just just like talk to it and like let it really try to figure out. And yes, you can still edit it and then do these kinds of things, but kind of like discourage a little of that and more just like let just talk to Claude to go figure it out. Um, the second thing was it was really trying to express that actually like code is a is a a way to solve for things that you wouldn't normally think would be the way. So a lot of people who have built kind of generative um you know like slide decks or designs or whatever um will pick uh the way of like they have like some kind of design system you integrate against design system. It's almost the traditional like classic wissywig style of designing something. And with like quad design, it was like okay, can we try to just like use code purely have Claude generate that code and would it like do a good job? And we found through some experiments early on. It's like actually it looks like it can kind of do that and how can we kind of showcase that uh to the world. So that's like an example. We have a lot of other internal projects and this kind of falls in the category of like expressing form factor. We'll all try it out internally. it'll be super cool for like two weeks and then we move on to the next thing. We never even ship the thing frankly. But yeah, we actually do a lot of product experimentation in that area and that's like our labs team. And then there's like the second category which is that we actually do look at TAM like we're a business. We do look at TAM. We do look at areas that we think uh you know there' be reasonable agentic like operations that would happen in those areas.
嗯,比如 Claude design 就有点这个意思。你要是换个角度眯着眼看,可能会觉得我们这是把设计当成一个垂直领域在切入。但更多时候,如果你看看我们在那个产品上到底想干什么,其实里面做了几个决定。第一个是,你其实可以把越来越多的事情丢给 Claude 去做。所以它有意做得比较「有主见」——就是你直接跟它说话,让它自己去把事情琢磨明白。当然你还是可以手动去编辑、去做那些操作,但它会有点刻意地弱化这一块,更多是让你就直接跟 Claude 聊、让它去搞定。第二件事是,它真正想表达的是:用代码其实可以去解决一些你平常根本不会觉得该这么解的问题。很多做生成式 PPT、生成式设计之类东西的人,选的路子都是搞一套 design system,然后你去对着这套 design system 集成——这几乎就是那种传统、经典的所见即所得(WYSIWYG)式的设计方式。而在 Claude design 里,我们想的是:能不能就纯粹用代码、让 Claude 直接生成那些代码,它能不能干得漂亮?我们早期做了些实验,发现——它好像还真能做到,那我们怎么把这个能力展示给全世界看。这就是一个例子。我们内部还有很多别的项目,也都属于「表达 form factor」这一类。我们都会在内部先试一把,火个两周,然后就转去做下一个了。说实话,很多东西我们压根就没上线过。但我们确实在这个方向上做了大量的产品实验,这就是我们的 labs 团队在干的事。然后还有第二类,就是我们确实会去看 TAM——我们毕竟是一门生意,我们确实会看 TAM,会去看那些我们觉得会有相当规模 agentic 操作发生的领域。
[20:07] Angela
Uh we do tend to have an orientation towards things that are more tokenheavy. And by token heavy or token hungry maybe is the way I would say that is like what we mean is like you know you for spending once you spend a like call it like one turn you look at the end of that turn and you say like am I done or am I actually so glad that I did that thing I want to do more of that thing we like industries where it's like the answer to that question you say I want to do more of that thing so coding is obviously the one that we all know and the great thing about coding is that what it's actually doing is that like once you finished a turn you look at that and you're like that was incredible I'm like unlocked I'm going to do like more. I'm going to build more. I can do more. And there's other services where it's like actually when you finish that turn, you completed the job and you just move on. You know what I mean? Um, and so we tend to like go into the ones that are a bit more like there's this kind of like iterative flow. You're going to build more, generate more together. Um, and then the last angle that we kind of take a look at is just sort of like, you know, there's going to be certain business functions that we're like, they are the buyer that we like to go to. We want to help them optimize their workflows, help them create better products there. And I think we've been pretty transparent with some of the verticalization. Like we've done like finance, we've done like legal um and we've tried to kind of like narrow on into specific areas where we feel like by having the right context and the right tools and putting it together in a good form factor is probably useful um for us to to be able to do.
我们确实偏向于那些更「吃 token」的场景。所谓吃 token、或者说 token 饥渴,我想表达的意思是:你花掉一轮(就叫一个 turn 吧)之后,你看着这一轮的结果,你会问自己——我是「做完了」,还是「我太庆幸我干了这件事、我还想再多干点」?我们喜欢的,是那种你对这个问题的回答是「我还想再多干点」的行业。编程显然就是我们都熟的那个例子。编程的妙处在于:你干完一轮,回头一看,会觉得「太牛了,我一下被解锁了,我要做更多、build 更多、我能干更多了」。而另一些服务则是:你干完那一轮,活儿就完事了,你就走人了,你懂我意思吧?所以我们倾向于切入那种更有「迭代流」的场景——你会一起 build 更多、生成更多。我们看的最后一个角度是:总有一些特定的业务职能,是我们觉得「他们就是我们想去服务的那个买家」。我们想帮他们优化工作流、帮他们在那里做出更好的产品。我们在垂直化这件事上一直挺透明的——比如我们做了金融、做了法律,我们一直想收窄聚焦到一些具体领域:在那些地方,只要有对的 context、对的工具,再用一个好的 form factor 把它们组装到一起,对我们来说多半就是有用的、能做成的。
[21:25] Angela
And in each of those areas, we do we're trying to do a bit of like showing the art of the possible across all the different ways that you would accomplish those outcomes. And so for you know like finance for example is a good one. Um, you know, we you could be a company that solves problems in finance and you could build directly on the messages API and you can just get some tokens and you can build everything else on top or you could be someone who builds on cloud managed agents. You can get a lot more out of the box or you could say I'm going to build a plug-in that or like a connector right that's going to sit within one of our products and within those form factors. When we did recently, we launched like claude for financial services is like, "Okay, cool. We've got packages of skills and things like this that you could choose to use within our product, within other people's products. We even launch like cookbooks on here's how you would use cloud managed agents to go and do these things." And so, I think for us, it's all kind of an experimentation around like, you know, we provide people all these different pieces and see kind of where they run with it. And then sometimes we put together products that are just packaging of all of these things like claw tag I think is a really good example like we had been seeing people in the industry go and say like Shopify did this with River um Square Block recently did this with Builderbot. Um, there's like a few of these examples where people said, "I'm going to pro I'm going to build like an agentic platform internal to my company and I'm going to try to give it all the right context and I'm going to make it accessible from Slack or from various other um, you know, platforms that you'd want it to be accessible at." And I think Claude tag was very much a packaging of all those same things that anybody could choose to build something similar but this is how we're kind of like well this is how we're doing it internally and if you would like to just kind of plug in and go here's what that looks like.
而且在这每一个领域里,我们都在试着做一点「展示可能性的艺术」——把达成那些结果的各种不同路径都秀出来。就拿金融来说,这是个好例子。你可以是一家专门解决金融问题的公司,直接建在 messages API 上,只管拿 token,其余全都自己在上面 build;你也可以选择建在 Claude managed agents 上,那样开箱即用能拿到的东西就多得多;或者你也可以说,我要做一个插件、或者说一个 connector,让它嵌在我们某个产品里、在那些 form factor 之内运行。我们最近上线了 Claude for Financial Services,思路就是:「好,我们准备好了一批打包好的 skills 之类的东西,你可以选择在我们的产品里用、也可以在别人的产品里用。」我们甚至还上线了 cookbooks,教你「该怎么用 Claude managed agents 去把这些事做出来」。所以对我们来说,这整个就是一种实验——我们把这些不同的零件都提供给大家,看他们能拿去跑出什么花样来。然后有时候,我们会把这些东西打包成产品,Claude Tag 就是个很好的例子。我们一直看到行业里有人这么干——比如 Shopify 用 River 做了这事,Square(Block)最近用 Builderbot 也做了。有那么几个这样的例子,大家都说:「我要在公司内部 build 一个 agentic 平台,尽量给它喂上所有对的 context,还要让它能从 Slack、或者从别的各种你想让它出现的平台上被调用。」而 Claude Tag 差不多就是把这同样一堆东西打了个包——任何人都可以选择自己去 build 一个类似的,但这就是我们内部的做法;如果你想直接接进来用,那它大概就长这样。
[23:11] Host
What do you think people misunderstood about cloud tag? Because there was all this like ruckus about oh my gosh it's just a slackbot like tell tell us [laughter] what the magic of tag is.
你觉得大家对 Claude Tag 最误解的地方是什么?因为当时闹得沸沸扬扬的——「天哪,这不就是个 Slack 机器人嘛」——跟我们说说(笑),Tag 真正神奇的地方到底在哪?
[23:19] Angela
No I think it's a great question. Um and I I do think it actually showcases a little bit of where maybe the future could be going. Um, yeah. I think like the I think if you look at products in the past, people are like, "Oh, you really attach to like the form or the the UI almost, right? Like it looks like this." So, it's like super cool. Um, and I think when you look at like tag, uh, it like yeah, like the way you interact with it is that you like literally tag it in Slack. Uh, and so yeah, that is like the interface, but that's not really the important part. The important part, um, is all the kind of like context engineering and like architecture that we put underneath the hood. So that tag just works. It really should just like just feel like a co-orker like a co, you know, if you go to a company and you onboard and a co-orker comes into your channel and then you can chat with it. It's proactive. It figured out like what's like useful you and um it just gets stuff like done for you. And so if you think about, you know, especially like nontechnical audiences, this is like it's a huge unlock. you just you literally create a channel and then you atclude or sometimes you don't even atclude and you're like hey I want to be able to do this and do that and I can't figure out this and how do I actually like submit an expense report again and traditionally you think about how to solve that workflow you are going all over the place and you're talking to your manager you're talking to your spin buddy and it's really really complicated and uh today now you just like go talk to cla tag and we do a lot of the hard work on doing the context engineering the proactivity a lot of the harness pieces I think Andre Kaparthi said it really well it's like it's like an org level harness There's a lot of like complexity baked into that like Kayla mentioned like you can use our APIs to go and construct that. You have to do a lot of the experimentation yourself obviously but this is like an opinionated take from anthropic on like how you can have this really awesome always on uh kind of agent for your entire entire company.
不,我觉得这是个特别好的问题。而且我确实觉得,它多少展示了未来可能的走向。我是这么想的:你看以前的那些产品,大家往往会特别执着于它的形态、或者说 UI 本身——「它长这样,所以超酷」。而你看 Tag,你跟它交互的方式,字面意义上就是在 Slack 里 @ 它一下——所以这确实是那个界面没错,但这其实并不是重点。重点是我们在引擎盖底下做的那一整套 context engineering 和架构,好让 Tag 就这么「直接能用」。它真的应该让你感觉就像个同事——你想想,你进一家公司、走完入职流程,一个同事进到你的频道里,你就能跟它聊天。它是主动的,它能搞明白什么东西对你有用,然后就直接帮你把事情办了。所以你想想,尤其是对那些非技术背景的用户,这简直是个巨大的解锁:你只要建一个频道,然后 @Claude,有时候你甚至都不用 @Claude,你就说「嘿,我想做这个、做那个,我搞不定这个,还有我到底该怎么再提交一次报销单啊?」——传统上你要解决这种工作流,你得到处乱窜,去问你的主管、问你的搭档,那真的特别特别麻烦。而现在,你直接去跟 Claude Tag 说就行,那些苦活累活——context engineering、主动性、还有一大堆 harness 的部分——都由我们来做。我觉得 Andrej Karpathy 有句话说得特别到位,他说这就像是一个「组织级别的 harness」。这里面烘焙进去了大量的复杂度,就像 Katelyn 提到的,你可以用我们的 API 去自己搭出这套东西——当然你自己得做大量的实验——但这是 Anthropic 给出的一个有主见的方案:你可以怎样为你整个公司搞出这么一个特别棒的、always-on 的 agent。
[25:07] Angela
And the bit that's like futuristic I guess is like a lot of that complexity is actually like it's like an iceberg. is like all the stuff underneath it that's actually becoming the harder and harder and like useful part that we're trying to like push through. And I think we'll see more and more like that kind of like tip bit that's like outside in the water. It's just like the interface can actually constantly swap like today, right? Like Slack is a place where a lot of people collaborate, a lot of business collaborate, but also a lot of people collaborate in teams and some people collaborate by a WhatsApp group um or they text each other or they may some people still email each other and like those could be the form factors that actually completely you can imagine agents just going there and being and they're almost taking up the same form factors as humans have taken up. It was almost like a very almost like boring take, but it's actually like I feel like the most like forward one because you want the agent and you want AI to basically be like another person and it's helping you, but it's like you know very intelligent can figure out all the context and you can always have it to be a really helpful assistant.
而我说的那个「有未来感」的地方在于:这里面很多复杂度其实就像一座冰山——真正越来越难、也越来越有价值的,是水面底下那一大堆东西,那才是我们在使劲往前推的部分。我觉得我们会越来越多地看到这种「露在水面外的那一小角」的情形:界面其实是可以不断替换的,就像今天,Slack 是很多人、很多企业协作的地方,但也有很多人在 Teams 里协作,有些人用 WhatsApp 群协作,有些人互相发短信,还有些人到现在还在用邮件往来——而这些都可能成为真正的 form factor:你完全可以想象 agent 就直接去到那些地方待着,几乎就占据了人类原本占据的那些形态。这几乎是个非常无聊、非常平淡的判断,但我反倒觉得它是最有前瞻性的一个——因为你就是想让 agent、想让 AI 基本上就像另一个人一样在帮着你,只不过它非常聪明、能搞清楚所有的 context,你随时都能让它当一个特别得力的助手。
[26:03] Host
Totally. You talked about context and then harnesses quite a bit and so your team is just, you know, has such an opinionated point of view on like what it takes to build an exceptional agent. I imagine a lot of that comes down to the context engineering and the harnesses.
完全同意。你刚才挺多地聊到了 context 和 harness,你们团队对于「打造一个卓越的 agent 到底需要什么」有一套特别有主见的观点。我猜这里面很大一部分,都归结到 context engineering 和 harness 上。
[26:16]
Totally.
完全正确。
[26:16] Host
Maybe like what best practices or advice would you would you share with people about what you need to get right on the harness and what you need to get right on the context.
那也许你们可以分享一些最佳实践、或者建议——关于在 harness 上你必须做对哪些事、在 context 上你又必须做对哪些事?
[26:24] Katelyn
Yeah, I think so. It's interesting because we've kind of talked about, you know, we launched cloud manage agents as this like very generic but high performing harness because we've done all the nitty-gritty work that's actually like really boring and not super interesting around how do you deal with prom caching? How do you deal with context management? You like clear old stuff out of the window. Sometimes you like call tools programmatically so you don't pull everything into the context window and you can keep it clean. There's a lot of those sort of details on the lower level harness layer. Um and I think honestly like best practices are just stuff like prom caching. Do it. You're going to save a lot of money and and token costs. Um, obviously like try to keep your context window clear and then putting those things together in uh a harness that will be performant is is you know sometimes specific to the task that you're trying to accomplish, right? And then of course evals. Um I'm surprised we got this far into this thing before one of us said the word evals, but like you need evals um to make sure that what you're trying to accomplish is performance. Um, but I think where we're starting to go, and Angela mentioned this a little bit earlier, is more of a concept of strategies or metah harnesses because I do think that yes, you can again make this lower level harness is going to be performant and maybe that's interesting for you to do yourself or maybe not and you offload it to us. But this concept that you can take any given token and spend that token on just executing or you could take that same token and choose to actually reflect on your past agentic sessions and write learnings to memory so that the next agent does a good job or you could take that token and advise with a bigger model so that a smaller model can execute and do a better job. Um or you can say execute execute and then like a greater comes in is like did you do a good job? No, you didn't try again. Right? And so I think the the like interesting innovation is going to come more at that higher level on like the meta level, right?
嗯,我想想。这挺有意思的,因为我们前面聊过,我们把 Claude managed agents 做成了一个非常通用、但性能很强的 harness,就是因为我们已经把那些又苦又无聊、一点都不性感的细活全干了——比如你怎么处理 prompt caching?怎么做 context 管理?你得把窗口里的旧东西清出去,有时候你得用编程的方式去调工具,这样就不会把所有东西全拉进 context window,好让它保持干净。在底层 harness 这一层,有一大堆这样的细节。老实说,最佳实践其实就是些这样的事:prompt caching,做就对了,你能省下一大笔钱、省下一大堆 token 成本。当然还有,尽量让你的 context window 保持清爽,然后把这些东西组装成一个高性能的 harness——这有时候要看你想完成的具体任务而定。当然还有 evals。我挺意外我们聊了这么久,才有人第一次说出 evals 这个词,但你确实需要 evals,来确保你想达成的东西是真的有性能保障的。不过我觉得,我们正在往前走的方向——Angela 前面稍微提过——更多是一种叫 strategies、或者说 meta-harness 的概念。因为我确实觉得,是的,你可以把这种底层 harness 做得很高性能,也许你自己去做这件事觉得挺有意思,也许你觉得没意思、那就把它甩给我们。但还有这么一个概念:你可以拿任何一个给定的 token,把它花在「纯执行」上;你也可以拿同样这个 token,选择让它去回顾你过去那些 agentic 会话、把学到的东西写进 memory,好让下一个 agent 干得更好;或者你可以拿这个 token 去跟一个更大的模型商量一下,这样一个更小的模型就能执行、并且干得更好;又或者你可以说,执行、执行,然后来个「裁判」进来问:你干得好不好?不好,你没做好,重来。所以我觉得,真正有意思的创新,会更多地发生在那个更高的、meta 的层面上。
[28:25] Katelyn
And I think optimizing within those strategies is something that our team is really excited about and we're starting to do a lot of work there. Um, and I think a lot of other people are starting to feel really excited about this concept of strategies and like the jobs you give to tokens because again like yes, there's best practices on stuff like your prom caching and exactly how you clear stuff out of your context window and how you write your evals and like a lot of things like this, but I I don't know that there's necessarily so much juice to squeeze in a lot of cases out of that layer as compared to a layer higher than that.
而在这些 strategies 内部去做优化,是我们团队特别兴奋的一件事,我们也开始在那儿投入大量的工作。我觉得很多别的人也开始对这个概念——strategies,以及「你到底给 token 派什么活」——感到特别兴奋。因为,还是那句话,在 prompt caching 该怎么做、context window 里的东西到底怎么清、evals 怎么写这些事情上,确实是有最佳实践的,诸如此类还有很多;但我不确定在很多情况下,那一层还有多少油水可榨——相比之下,比它更高的那一层要划算得多。
[28:53] Angela
Yeah. And one of the reasons for that I think is it has to do with the generations of the the models. I if you look like two years ago, a lot of the harness was like a scaffold to kind of like tell the model to go from point A to point B. And you had to like you really had to like build in a lot. You practically build one wall here and one wall here. So like the thing would go in a straight line. And now the models are actually very very steerable. Um and so a lot of that steering you could just put in the prompt, right? Like go do go from point A to point B and the model like will go from point A to point B. So, a lot of if you have harnesses um that are like designed to kind of do that kind of like steering, you can delete that part. Like that part we actually frequently encourage like you can delete part of those harnesses. I think various people have said things along those lines. And that's I think what people often times mean when they're like the model will kind of consume some of the scaffolding and like in that sense like for sure if your scaffolding is telling it to go in direction um that it can just intelligently figure out like that I think will increasingly continue to to be so. But as a result of of this, what the harness needs to start doing is more allow it to run longer. And so that's where like that execution bit tends to be. I think like it sounds like a maybe somewhat silly point, but I do think it results in a lot of differences because because you can go in the direction that you tell it to go. You obviously don't want it to stop at B. You're going to be like, "Okay, now go from B to C and then go to F and then go to Z and then come back to me on A." You know, something funky like that. In order to be able to do a lot of those things, the kinds of harnesses that you do are less the steering harness and it's more like these kind of strategy harnesses that Caitlyn's mentioning, which allows you to operate at a slightly higher level of thinking which matches I think a lot of the intelligence gains that we're trying to see with the model.
对。而我觉得其中一个原因,跟模型的「代际」有关。你回头看两年前,很多 harness 其实就是一层脚手架,用来把模型从 A 点「引导」到 B 点。你真的得往里塞很多东西,你几乎是要在这儿砌一堵墙、在那儿砌一堵墙,好让这东西能走一条直线。而现在的模型其实已经非常非常可引导(steerable)了,所以很多那种引导,你直接写进 prompt 里就行——你就说「去,从 A 走到 B」,模型就真的会从 A 走到 B。所以,如果你的 harness 里有一部分是专门为了做这种「引导」而设计的,那部分你其实是可以删掉的。我们其实经常鼓励大家:这些 harness 里的那部分,你可以删掉。我觉得也有不少人说过类似的话。这也就是大家常说的「模型会把一部分脚手架给吃掉」的意思——从这个角度看,如果你的脚手架是在告诉它往某个方向走、而这个方向它自己就能聪明地想明白,那我觉得这种情况只会越来越普遍。但正因如此,harness 接下来需要开始做的,更多是让它能跑得更久。这也就是那个「执行」环节所在的地方。我知道这听上去可能有点像个挺傻的点,但我确实觉得它会带来很多不一样的地方——因为既然它能朝你指的方向走,你显然不会想让它到 B 就停下。你会说:「好,现在从 B 走到 C,然后去 F,再去 Z,然后就 A 这件事回来找我。」诸如此类有点花的操作。而要能做到这一大堆事情,你要用的那类 harness,就不再是「引导型」的 harness,而更像是 Katelyn 提到的那种「strategy harness」——它让你能在一个稍微更高的思维层级上去操作,我觉得这正好跟我们想在模型上看到的那些智能提升相匹配。
[30:28] Host
Do you think task specific harnesses make sense or a vertical specific or task specific harnesses?
你觉得针对特定任务的 harness 是有意义的吗?或者说,针对特定垂直领域、特定任务去做 harness?
[30:33] Angela
I think people have different opinions on this. Like our opinion is yes. I don't think there's like a general harness. I think there are some capabilities that are obviously very general and they tend to like be very useful. Uh like coding is a capability that like is very useful because you can use it across so many things and software as uh you know just like eaten so much of of what is capable. So our ability to like write software is therefore useful. I think when you think about like very very specific types of domains they're going to require like a couple of pieces of the harness to be sort of like customized. One of that uh I do think is how you choose to kind of like handle sort of like errors uh between when you do something and you hand something off to the model. Um, so in like domains where you require like an extreme level of verification, that logic of how you handle that ver like it again, I think it sounds small, but like I totally understand why some people feel like they really want to own the harness because tweaking that last bit will give you a ton of juice and especially domains like like legal and finance where there's a lot of consequences um to you not getting it perfectly correct like is really going to matter and that's going to be the difference between your product and someone else's product being the thing that the user ultimately uses. Um and then there are other domains for which like I would say uh it's not going to matter as much because you're able to compress it into like a general model capability. So the tweaks that I guess like you know where we feel like the domain specificity is really going to matter is the specific like verification logic between the model and your execution. And then um I think it's going to be about like some of these kind of like higher order strategies on how well um you're able to actually like allocate your token budget. Um, I think the context bit is actually a little like overdone.
我觉得大家在这件事上看法不一。我们的观点是:有意义。我不认为存在一个通用的 harness。有些能力显然非常通用,也往往特别有用——比如 coding 就是一种极其有用的能力,因为它能被用在特别多的场景里,而软件本身已经吞掉了这么多原本需要人做的事,所以我们写软件的能力也就跟着变得有用。但当你面对非常非常具体的领域时,harness 里总有那么几个部件是需要定制的。其中一个我确实觉得很关键,就是你怎么去处理错误——也就是当你做完一件事、把某个东西交给模型的时候,这中间的错误怎么 handle。所以在那些对验证要求极高的领域里,这套『如何处理验证』的逻辑……听起来是件小事,但我完全能理解为什么有些人特别想自己掌控 harness,因为把最后这一点点调好,能榨出巨大的收益。尤其在 legal、finance 这类领域,一旦你没做到完全正确,后果很严重,这会真正决定用户最终到底用你的产品还是别人的产品。而另一些领域里,我会说这就没那么要紧了,因为你能把它压缩进一个通用的模型能力里。所以我们觉得领域特异性真正重要的地方,是模型和你的执行之间那套具体的验证逻辑;再就是一些更高阶的 strategies——你能多好地去分配你的 token budget。至于 context 这块,我其实觉得有点被夸大了。
[32:20] Angela
Like yes, you're going to like throw in context and like that's uh but any harness can actually handle a lot of context and so that's just more like you have the data and if you have the data then obviously you're you're uniquely qualified to do something useful.
当然,你会往里塞 context,这……不过其实任何 harness 都能处理大量 context,所以这更多只是说明你手上有数据;而只要你有数据,你显然就有独特的资格去做点有用的事。
[32:24] Katelyn
Yeah. And I think when people say harnesses they often mean a lot of different things and I think this is why in part there's so many different opinions on this. Like you can think of a harness as literally just like a loop um that's like okay cool like user model user model tool you know like that sort of thing. Um then you could think of the harness as also all of the tools that are packaged up with the harness right and and there's just like a lot of different definitions of these things. And I think the stuff that can be pretty generic and like less interesting to own and and deal with is what I was kind of saying earlier is like getting your prompt caching right right like maybe that is not the world's most interesting thing. Choosing to clear out old tool calls from the context window and and things like that right are like maybe a little bit less interesting and you like go a layer higher into some of the stuff Angela's talking about and then you get into like okay yeah these are things that I might want to own and control. And so it's interesting with cloud manage agents like the thing that we built today, we call it higher order, but it's not really like that high order in the sense that you can choose to define all of the tools that you want to bring in as custom tools with the harness, right? And like we give you a lot of knobs to control, you can define skills, you can do your system prompts, you can do a whole bunch of different things, MCP servers and things like this. And I think where you know we want to get to is a point where you can literally just tell an agent here's the outcome I want and here's the budget that I want to spend like ready set go and you may be like don't think about any of those things underneath. And so I think there's just a few different layers of this right that for certain things like you might want to sit at a different layer of what you actually go and control. Um and you can probably get better outcomes within some of those layers by doing a little bit more optimization work.
对。而且我觉得,大家说 harness 的时候,往往指的是很不一样的东西,这也是为什么这件事上会有这么多不同观点。你可以把 harness 理解成就是一个循环——好,user、model、user、model、tool,就这么个东西;你也可以把 harness 理解成连同它一起打包的所有 tools。这些说法差别很大。那些相当通用、拥有和打理起来没那么有意思的部分,就是我前面说的——比如把你的 prompt caching 弄对,这可能不是世界上最有意思的事;再比如选择把旧的 tool call 从 context window 里清掉,诸如此类,这些可能也没那么有意思。而你往上再走一层,进到 Angela 说的那些东西,你就会觉得:对,这些才是我可能想要自己掌控的。所以 Claude managed agents(也就是我们今天做的这个东西)挺有意思的,我们叫它 higher order,但其实也没那么『高阶』,因为你可以把想引入的所有 tools 都定义成 harness 里的 custom tools。我们给你很多旋钮去控制:你能定义 skills、写你的 system prompt、做一大堆别的事,MCP server 之类的都行。而我们想要到达的那个点,是你真的可以直接告诉一个 agent:这是我想要的结果,这是我愿意花的预算,预备——开始;至于底下那些细节,你可能压根不用去想。所以这里其实有好几个不同的层,对某些事情你可能想坐在不同的那一层去真正控制;而且在其中某些层里,多做一点优化工作,你大概能拿到更好的结果。
[34:07] Host
Very cool. One of the things I'm curious about and one that I love about infrastructure and platform teams is that you get to see what the most advanced users in the world are using and learn from them. I'm curious what are some things that you're seeing and learning from from the people building on your platform.
很酷。我特别好奇的一点——也是我特别喜欢 infra 和平台团队的地方——就是你们能看到全世界最先进的用户在用什么,并从他们身上学习。我很好奇,你们从在你们平台上做开发的人身上,看到和学到了哪些东西?
[34:21] Angela
There's some people that have been doing some really funky ways of like handling context. Um we ourselves explore this a lot. That's actually like one of the reasons why TAG is like uh such a great product is like there's a lot of really awesome like context kind of engineering that that's happening. Um, we've seen some teams be really clever about like how they do that and they are able to kind of think through like, okay, if I have all these contacts in a bunch of different places, how can I proactively go reach out to them? How can I try to generate enough like um permissions across each of them? So, and then feed that all into like an agent. And it's interesting that like um I guess like this is kind of the level of innovation that like we're actually like very excited by. It doesn't express itself as like a completely different product form factor. Um, but what it actually does express itself as is like maximally useful to users and we've been seeing this more and more with like inter actually like internal use cases instead of like external ones. So like companies who are becoming more AI native basically they're the ones we're seeing increasingly more and more innovation out of and so you know we've had like customers try to do this for their like they've built their own like custom SDLC kind of setup in very very innovative ways. We've had uh ones who do that for like their entire back office and just like the kind of nuances of how they like stream in context I think has been like actually really interesting in terms of like how they've been putting together the pieces. So that's been like one category that's been like really really like fascinating. Uh another category that's been like really interesting has actually been with companies that are dealing with like really old school software. And so there's a lot of like healthcare companies um that we kind of engage with and you know like they're like the the systems I'm working with they don't even have APIs. like that's that's a a dream. Um and so you know how can they use computer use uh and things like this to be able to start to kind of automate and create more connectivity with our systems.
有些人处理 context 的方式非常新奇。我们自己也在大量探索这块。这其实也是为什么 TAG 是个这么棒的产品——里面有很多特别精彩的 context engineering。我们看到有些团队在这方面特别聪明,他们能想清楚:好,如果我的 context 散落在一堆不同的地方,我怎么主动地去把它们够到?我怎么在每一处都生成足够的权限?然后把这一切都喂进一个 agent。有意思的是——我想这大概就是让我们真正兴奋的那种创新层次——它并不表现为一个完全不同的产品形态,而是表现为对用户极其有用。而且我们越来越多地是在内部用例里、而不是外部用例里看到这一点。也就是说,那些正变得更 AI native 的公司,才是我们看到创新越来越多的地方。所以我们有客户用非常非常创新的方式,为自己搭了一套定制的 SDLC;也有客户把整个后台办公室都这么搞了。他们把 context 流式接入的那些细微处理方式,我觉得从『怎么把这些拼图拼起来』的角度看,真的特别有意思。这是让我特别着迷的一类。另一类特别有意思的,其实是那些在跟非常老派软件打交道的公司。有很多 healthcare 公司我们会打交道,他们会说:我在用的这些系统,连 API 都没有——有 API 简直是奢望。所以他们怎么能用上 computer use 之类的手段,开始去自动化、并跟我们的系统建立更多连接。
[36:06] Angela
Um and that area of innovation I think has been really exciting. It's been really interesting to see people try all sorts of crazy stuff from like taking a laptop and trying to like run a bunch of things on it to autogenerate a bunch of things that then their agents can go and use. Um, and this has actually been probably like an area of um, I think a lot of innovation coming from a lot of our customers that we want to find ways to like support better and see like okay maybe there are like how can we make this easier for you? How can we help you with some standardization? How can we get it so that you know like you can just have a spec and then claude can then respect it and so it's much easier for you to organically connect a lot of these things. But yeah, maybe the the general theme I would just give you is like interestingly a lot of the innovation that's most exciting out there right now has been uh this kind of like context and connectivity layer which has been really fascinating.
这个创新领域我觉得真的特别让人兴奋。看着大家去尝试各种疯狂玩法特别有意思——比如拿一台笔记本电脑,试着在上面跑一堆东西,自动生成一批东西,然后他们的 agent 就能拿去用。这块其实很可能是——我觉得我们很多客户都在这里做出大量创新,而我们想找办法把它支持得更好,去想:好,也许有一些办法能让这件事对你更容易?我们怎么能帮你做一些标准化?我们怎么能做到——比如你只要有一份 spec,Claude 就能遵循它,这样你就能非常自然地把这一大堆东西连起来。总之,也许我给你的一个总的主题就是:有意思的是,眼下最让人兴奋的很多创新,都来自这种 context 和连接层,这真的特别迷人。
[36:49] Katelyn
Yeah. Like a good one in that um we were working with a customer who they've built some agents on cloud manage agents. They also have some agents they built on other models and other platforms and they've kind of optimized each of these agents to be good at the things that they want. They want these agents to all be able to work well together. Um, and they kind of were like, "Wow, Galaxy brain. Like, what if I expose an MCP server on top of this agent so that it can then go and like have this other agent call a tool on that agent, right? And and have these things just be more modular and be able to work together." And we were like, "Yeah, totally." And we sat down with them and worked through it and and it worked perfectly and it was pretty cool. And so, we're seeing a lot of again that connectivity layer that I think is one of the cooler areas where people are innovating. But outside of that, one thing that has been cool is just seeing the shift in I guess like industry trends of where we're seeing a lot of our usage come from like talked a lot about coding like coding as a category like of course absolutely explode in. There's so much going on there and we're starting to see some of these emerging trends like more recently. Um, we're starting to see manufacturing really pick up as just a category where people are building with AI and like one of our PMs like getting on a flight to Detroit to go like figure out what these customers like what they need and what's going on. And so I think we're going to start to see a lot more just kind of like outside of the box of what people think about today sort of use cases which we're really excited about. H
对。有个很好的例子:我们跟一个客户合作,他们在 Claude managed agents 上搭了一些 agent,同时也在其他模型、其他平台上搭了一些 agent,并且把每个 agent 都调优到擅长他们想要的那些事。他们希望这些 agent 都能很好地协同工作。于是他们就想:『哇,galaxy brain(脑洞大开)——如果我在这个 agent 上头暴露一个 MCP server,让另一个 agent 能去调用这个 agent 上的一个 tool,会怎么样?这样这些东西就能更模块化、能协同工作。』我们说:『对,完全可以。』然后我们就跟他们坐下来一起搞,结果跑得完美,挺酷的。所以我们又一次看到那个连接层,我觉得这是大家在创新的更酷的领域之一。除此之外,还有一件很酷的事,就是看到行业趋势的转变——我们大量使用量的来源正在变化。前面聊了很多 coding,coding 作为一个品类当然是绝对爆发式增长,那里发生的事情太多了。而我们最近开始看到一些新兴趋势,比如 manufacturing(制造业)真的开始起来,成为一个大家用 AI 来做开发的品类。我们有个 PM 都直接飞去底特律,去搞清楚这些客户到底需要什么、都在发生什么。所以我觉得,我们会开始看到很多超出今天大家惯常想象的用例,这让我们非常兴奋。
[38:13] Host
it seems like there's now there's a we went through a token maxing moment of history and now there's like the token rationalization moments of history. [laughter] What are your thoughts on that and like what what should companies be doing and then how how does the platform team think about uh enabling that?
看起来,我们现在……我们经历了历史上一个『token maxing(把 token 用到极致)』的时刻,而现在则进入了历史上『token rationalization(把 token 用得合理)』的时刻。[笑] 你们对此怎么看?公司应该怎么做?平台团队又是怎么考虑去支持这件事的?
[38:28] Angela
Yeah, I mean it it makes sense. Uh it it makes sense from the high you start to rationalize. I I really like that framing and I think there's like a couple things that that are like top of mind for us on this front. I think like again it makes sense and as these models get more and more capable you're going to hit like levels of intelligence max maxing that are like there that then you want to do the next kind of dimension and the next dimension after intelligence will either be cost or it will be speed. Um and you just kind of you know go through that across all possible tax complexities in the distribution. Um, and as we kind of see that like happen, you know, something that's like really top of mind for us that we kind of try to spend some time with users on is like what you don't want to do is like stop AI usage, right? Like that's kind of the wrong move. And we do actually see some of our our customers do that. So oftent times the way that AI spend has erupted inside their company has been through some kind of like uh shadow IT, you know, like their employees just like want to use it, they find a way, they end up procuring it themselves, and before you know it, like half your or has like found some way to have installed cloud code. And in that world it is kind of hard to to manage because these things are again like they're very token hungry ultimately. And so what we try to kind of encourage our customers is like okay you don't want to like stop the innovation like if you are getting returns on top of this you are shipping faster than ever before you can like run more operationally like uh efficient then those are gains. And so the area that we actually try to encourage people is like if there is a way for you to kind of construct again like a strategy that allows you to design an architecture that says like given a task assesses level of complexity.
对,这说得通。从……你开始变得理性,这是说得通的。我特别喜欢这个说法。在这件事上,有几点是我们特别在意的。我觉得——再说一遍,这说得通——随着这些模型越来越强,你会撞到一种『intelligence maxing(把智能用到极致)』的水平,到那时你就想去做下一个维度;而智能之后的下一个维度,要么是成本,要么是速度。你就这样,把分布里所有可能的任务复杂度都走一遍。当我们看着这件事发生的时候,我们特别在意、也会花时间跟用户聊的一点是:你最不该做的,就是叫停 AI 的使用——那基本是个错误的举动。而我们确实看到有些客户这么干。很多时候,AI 花费在一家公司内部爆发式增长,是通过某种 shadow IT(影子 IT)——员工就是想用,他们总能找到办法,最后自己把它采购下来;等你反应过来,半个组织已经不知怎么就装上了 Claude Code。在那种情况下确实挺难管的,因为这些东西说到底非常『吃 token』。所以我们试着鼓励客户的是:好,你不想叫停创新——如果你在这上面拿到了回报,你出货比以往任何时候都快,你能把运营跑得更高效,那这些都是收益。所以我们真正会鼓励大家去做的方向是:如果你有办法搭一套 strategy,让你能设计出一种架构,针对一个进来的任务去评估它的复杂度……
[40:06] Angela
I mean I'm effectively describing a router but like there are ways to do this that are like I think a bit better now and so like this task comes in has a certain level of complexity for that level of complexity like you can define some rules but for the most part right if it's like a hard task you should probably route that to like a big super smart model and if it's not a hard task you can route that to like cheaper models um designing that I think has a little bit of like there's a lot of technical complexity in that but it's like very very doable and we actually like encourage people to try those kinds of things I think ultimately
我这其实就是在描述一个 router,但现在有些做法我觉得会更好一点。所以,这个任务进来,带着某种复杂度;针对那个复杂度,你可以定义一些规则,但大体上——如果是个难任务,你大概应该把它路由到一个又大又超聪明的模型;如果不是难任务,你就可以路由到更便宜的模型。设计这套东西,里面确实有不少技术复杂度,但它非常非常可行,我们其实很鼓励大家去尝试这类做法。我觉得归根到底……
[40:20]
offer rather
……是去 offer(提供),而不是说——
[40:21] Angela
I I think within the clawed space it will like make sense. It's actually one of the strategies we imagine like designing because the way that we kind of thinking a lot of these things is like it almost feels like every month there was a new era of something. Um and if we take a step back like okay and this seems to be like really fast and so what are the different ways that are recomposable so we can redesign very quickly for any new whatever the cool thing is that month kind of like bit. Um and so this is like in that category of things where we feel like we can actually just like recompose a lot of our primitives and then design it. I think the bit that we do feel really strongly about on the model routing front is like we are designing our platform for Claude and we want to make sure that Claude is great at like solving all these things. So we'll like restrict to that space um rather than you know I don't think we're that interested in saying like okay and then you know you should route to a different model or whatever.
我觉得在 Claude 这个范围内,它是说得通的。这其实也是我们设想要去设计的 strategies 之一,因为我们思考这些事情的方式是:几乎感觉每个月都会冒出某个东西的新纪元。退一步看:好,这个节奏似乎真的很快,那有哪些可重新组合(recomposable)的方式,能让我们非常快地为任何新东西——不管那个月冒出来的酷东西是什么——重新设计一遍。所以这就属于那类『我们其实可以直接把我们的很多 primitives 重新组合、然后设计出来』的事。在 model routing 这件事上,我们确实有一点想得很明确:我们是在为 Claude 设计我们的平台,我们想确保 Claude 在解决所有这些问题上都很出色。所以我们会把范围限定在这个空间里,而不是——我不觉得我们有多大兴趣去说『好,然后你应该路由到另一个模型』之类的。
[41:11] Host
Makes sense.
有道理。
[41:11] Katelyn
Yeah. and and well some of that too is just like I think we have a strong belief that harnesses and and just like the agentic layer should be tuned to the model family that you use it with. And so I think there was a period where people were kind of like yeah cool I can like build a harness and build an agent and then just like plug in a different model underneath and they were excited about routers from that perspective. And I think we started to see um like Verscell just did this with harness agent for example like some of these players in the space like come up a layer of abstraction and say actually like plug in the whole harness and the whole agent that's tied to a model family which makes a lot of sense and so what we could provide is a little bit better smarter like how do you mix and match the right models within the model family underneath that thing if that makes sense. But yeah, on the general question of token maxing costs and these sorts of things, I think we're just kind of going through what feels like a normal natural cycle for companies and figuring out how to make the best use of this technology and run their businesses really well and really effectively. And um it's interesting like before working at Anthropic I was at Stripe and we were kind of in the very reasonable era of like we paid a lot of attention to our AWS bill and so you know if someone were to have built some background job and they like didn't quite configure it correctly and this thing's like burning through like CPU or whatever it is right like at any given moment and causing you know big increase in spend that's not actually worth it right like we have put in place the guardrails to find that and then go ask that engineer very nicely to please turn off their background job that's not like within the bounds of of what they should be spending for the thing they're trying to accomplish. I think those are the things with AI that people are going to start to go and figure out. And I think to Angela's point, a thing that gets dangerous is when you're kind of just like here's a cap and you're stuck within your cap like ready, set go.
对。而且这里面还有一部分是——我觉得我们有一个很强的信念:harness 以及整个 agentic 层,应该针对你所搭配使用的那个模型家族来调优。我记得有一段时间,大家的想法是:『好,酷,我可以搭一个 harness、搭一个 agent,然后在底下随便插一个不同的模型进去』,他们正是从这个角度对 router 感到兴奋。而我们开始看到——比如 Vercel 最近就用他们的 harness agent 做了这件事——这个领域里有些玩家往上抽象了一层,说:其实你应该把整个 harness、整个 agent 一起插进来,而它是绑定到某个模型家族的,这非常有道理。那么我们能提供的,就是在那个东西底下、在同一个模型家族内,怎么把合适的模型更好、更聪明地搭配组合起来——如果你懂我的意思。不过回到 token maxing、成本这类大问题上,我觉得我们只是在经历一个对公司来说挺正常、挺自然的周期,去摸索怎么把这项技术用到最好、怎么把生意跑得又好又高效。有意思的是,在来 Anthropic 之前我在 Stripe,我们当时处在一个非常理性的年代:我们特别关注 AWS 账单——如果有人搭了个后台 job 却没配好,这东西随时在疯狂烧 CPU 什么的,造成花费大涨,而这其实并不值得,那我们就已经建好了护栏(guardrail)去发现它,然后非常客气地去请那位工程师把这个后台 job 关掉,因为它超出了『为达成目标该花的范围』。我觉得 AI 上也会是这些事,大家会开始去摸索。而且照 Angela 说的,危险的地方在于:你只是简单地『给你一个上限,你就卡在这个上限里,预备——开始』。
[43:02] Katelyn
But I do think that encouraging innovation, encouraging people to, you know, create really excellent outcomes with this stuff and then coming in from the side and looking and saying like, okay, well, there are few different ways that we probably could have accomplished that outcome, right? And one is like you take Opus and you run it all night and you do something crazy. And another is maybe to get a little bit smarter with the strategies that you put together in order to create that same outcome within a lower cost. And I think that's the like next layer of thinking that everyone's going to start to do.
但我确实觉得,鼓励创新、鼓励大家用这些东西做出真正出色的成果,然后再从旁边切进来看一眼、说:好,其实我们大概有几种不同的方式能达成同样的成果——一种是你拿 Opus 通宵跑,搞出点疯狂的东西;另一种也许是在你为达成同样成果而组合起来的 strategies 上,变得更聪明一点,从而用更低的成本做到。我觉得这就是下一层的思考,是每个人都会开始去做的。
[43:31] Host
Very cool. Is there anything that you guys are excited about building over the next few months that you can share a hint at what might come next?
非常酷。接下来几个月里,有没有什么你们特别期待去做的东西,可以稍微透露一点接下来会有什么?
[43:37] Angela
Uh yeah. I mean I know we said this word like 20 million times. I apologize but like we really are trying to build ways for you to compose strategies. Um and so uh that is an area that that we're like trying to move into that kind of like yeah uh coordination layer of the abstraction. Um, and we want to start at this front because the types of problems that we see people building, they are at a layer where it's like in order to get the most return on this, you have to be a little clever about like what is the nature of the problem that you're solving. So to give you something like concrete like when you try to solve for like let's say you want to build an agent that's like trying to um do bug hunting and you could just send one off to go and do that and it's going to give you a certain type of return a level of return of possibility um and then people kind of get stuck at that and they're like okay my next options are I can like make a bigger I can just like swap the model for a different I probably bigger model um or I could like let it run like longer and that's pretty much like the only two like levers that you have to like try to make this like bug hunting agent. From a lot of experimentation, when we do these kinds of things, there's like actually the thing like those two those two things are still true, but you actually have like a third lever and tends to actually do a lot more than you think it does, which is that actually if you were to like best of end the thing, it would like give you a lot more returns. But like just to be just saying those words are fine and there's plenty of papers and people have published it to actually build that thing and put it into production so you can actually test it on users uh and see the results for yourself, that's like really really freaking hard. and you end up building all these like custom harnesses so on so forth like you know all that stuff.
嗯,有。我知道这个词我们已经说了大概两千万遍了,抱歉,但我们真的在努力打造让你能『组合 strategies(compose strategies)』的方式。这就是我们正试着切入的一个领域——那种协调层(coordination layer)的抽象。我们想从这个方向切入,是因为我们看到大家在搭建的那类问题,都处在这样一个层次:要想从中拿到最大回报,你得对『你要解决的问题本质是什么』稍微聪明一点。给你举个具体的例子:当你想解决——比如说你想搭一个 agent 去做 bug hunting(找 bug),你可以就派一个出去干这事,它会给你某种类型的回报、某种程度的可能性回报;然后大家往往就卡在这儿了,他们会想:好,我接下来的选项是,要么把模型换成一个不同的、大概更大的模型,要么让它跑得更久一点——差不多这就是你手上仅有的两个杠杆,去试着让这个 bug hunting agent 更强。但从大量实验来看,那两件事其实仍然成立,可你其实还有第三个杠杆,而且它往往比你以为的要管用得多:如果你对这件事做 best of N(跑 N 次取其中最好的一次),它会给你多得多的回报。不过话说回来——把这几个词说出来很轻松,也有一堆论文、有人已经发表过了,但真要把那个东西搭出来、放进生产环境,让你能在真实用户身上测试、亲眼看到结果,那可真是难得要命。你最后会搭出一堆定制的 harness,等等一大堆东西。
[45:17] Angela
Um but we're seeing like this is where the alpha is and it's hard and so like in the same very simple philosophy that we talked about at the beginning like if it's like gives you the return that you want and it's hard we're going to try to make it easy for you so then you can use it to then run the experiments you actually need to run.
但我们看到的是,真正的 alpha 就藏在这儿,而且它很难做。所以还是我们开头讲的那套很简单的理念——如果一件事能带给你想要的回报、但它又很难,那我们就想办法帮你把它变简单,这样你就能用它去跑那些你真正需要跑的实验。
[45:20] Host
Reminds me of when people are talking about agent swarms a year ago. It's some version of that.
这让我想起一年前大家聊 agent swarm(智能体集群)的时候,感觉就是那个东西的某种版本。
[45:25]
Has it been a whole year?
都已经过去整整一年了吗?
[45:26]
Yeah. Oh my god,
是啊。我的天。
[45:27]
I know. We're finally there.
可不是嘛,我们总算走到这一步了。
[45:28] Angela
Yes. Um yeah. No, I think that that's like that's a type of strategy. Exactly. In the same way that you have like, you know, one big one that separates a bunch as another type of strategy. And I think people have thought about this maybe the in the way of like human organization.
对。嗯,是的。我觉得那其实就是一种 strategy。没错。就像你也可以搞一个大的 agent、再从它分出一堆小的——那是另一种 strategy。我觉得大家过去可能一直是照着人类组织的方式来想这件事的。
[45:41] Angela
I guess it could be similar, but if you take it to kind of its ends, it's actually more just like the token has a job. And I think it's this job piece that we're we're really indexed on and um we see a lot of returns too and that's the thing that we want to spend time with users and the rest of the ecosystem on on like how can we just make that easier for folks to then experiment like we can give you like five jobs off the top of our head and we'll probably like that's what we have internally. Um and if we give this out to the rest of the ecosystem there's probably going to be like 100,000 200,000 who knows what other combinations that people could put together. Yeah, we want to be able to keep doing this hill climbing on like how do you get the most value, the most intelligence per dollar and just put that power in people's hands, but around the edges of that, we have these personas that have kind of just like things they have to work through in order to be able to like really deploy AI either within their companies or within their products. And um that's like the sort of enterprise ready security and compliance controls and things like this. But really even just like making the platform more modular in the right ways like being able to plug in different pieces of the solutions that we're building like I want to use memory for this thing over here, right? or whatever else it is and having a truly excellent developer experience around that because we spent a lot of time with enterprises who are like okay I have this like walled garden and I need to figure out exactly how I can plug these solutions in and so we're we've got a part of our team that's innovating on things like strategies and jobs and trying to help you maximize intelligence and they're like that's really cool but I can't actually use any of that for XYZ reasons. So I think solving those problems is really really important to us. But then the other persona is you know the like weekend developer who's like I want to go and build something useful for myself right and they're often doing that on top of our platform and on top of many other just pieces of developer platforms in the community.
它们可能确实有点像,但如果你把它推到极致,其实更像是「每个 token 都有一份自己的 job(活儿)」。我们真正重仓押注的就是这个 job 的概念,而且我们也看到了很多回报。这正是我们想花时间和用户、和整个生态一起去打磨的地方——怎么才能让大家更容易地去做实验。我们随口就能给你举出五个 job,这差不多也就是我们内部在用的那几个。但如果把这套东西开放给整个生态,人们能拼出来的组合估计得有十万、二十万种,谁知道还有多少呢。是的,我们想持续做这种 hill climbing(爬坡式优化)——怎么用每一块钱榨出最多的价值、最多的智能,然后把这份能力交到大家手里。但在这件事的周边,还有几类不同的 persona(用户画像),他们各自都有一些必须先迈过去的坎,才能真正在自己公司里、或者自己产品里把 AI 部署起来。比如那些面向企业的安全、合规管控之类的东西。但其实哪怕只是把平台按正确的方式做得更模块化一点——让你能把我们造的方案里的不同零件插拔进来,比如「我这块想用 memory」,或者别的什么——再围绕这一点提供真正一流的 developer experience,就已经很关键了。因为我们花了大量时间跟企业客户打交道,他们常常是这样:我这儿有个 walled garden(围墙花园),我得搞清楚到底怎么才能把这些方案接进来。所以我们团队里有一部分人在 strategies、jobs 这类东西上做创新,想帮你把智能最大化,可客户会说「这真的很酷,但因为这样那样的原因,我根本用不上」。所以在我们看来,解决这些问题真的特别特别重要。而另一类 persona,就是那种周末开发者——「我就想给自己搭个真正有用的东西」,他们往往就是在我们的平台之上、以及社区里许许多多别的开发者平台之上去搭的。
[47:26] Angela
And I think for some of those folks, there's more that we can do to be provide solutions that are maybe more open or more hackable or whatever it might be for those folks to kind of just like go wild with what we can offer them and have this really excellent developer experience. And so I think there's a lot of stuff that maybe I would put in the category of table stakes that I'm really excited about because I think those are the things that then unlock getting people to say, "Okay, yes, this thing works for me and now I can plug in on some of the stuff that you guys are doing." that's really innovative and hill climby to get more intelligence and save costs and things like that.
我觉得对其中一部分人来说,我们还能做更多——提供一些更开放、更可 hack(可自由折腾)的方案,让他们能拿着我们给的东西尽情发挥,同时享受一流的 developer experience。所以有一大堆东西,我可能会把它们归到「table stakes(入场门槛)」这一类,但我对它们真的很兴奋。因为正是这些东西才能起到解锁的作用——让人们说「好,这玩意儿对我确实管用,现在我可以接上你们正在做的那些真正有创新、能爬坡拿到更多智能、还能省成本的东西了」。
[48:01] Host
Wonderful. Caitlyn, Angela, I feel I mean you are building one of the most important developer platforms in the world and talking to the two of you over time. I just feel really optimistic that that platform is in very thoughtful uh hands that that care about the ecosystem. So, thank you for taking the time today to share what you're up to and um we look forward to what's ahead.
太好了。Katelyn、Angela,我是说,你们正在打造全世界最重要的开发者平台之一。跟你们俩这一路聊下来,我真心觉得很乐观——这个平台掌握在一双非常有想法、又真心在乎整个生态的手里。所以,谢谢你们今天抽时间来分享你们正在做的事,我们很期待接下来的进展。
[48:21]
Thanks for having us.
谢谢你们的邀请。
[48:22]
Thank you guys.
谢谢你们。
[48:33]
[music] [music]
[音乐]