ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.69 · 全文

Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic

频道: AI Engineer
视频: https://www.youtube.com/watch?v=uU5Gv2h8-9g
原文语言: en
统计: 共 115 轮 · Simon 46 · Cat 20 · Thariq 28 · Simon / Cat 2 · Cat / Thariq 1 · Cat / Simon 1


[0:12] Simon

Welcome to this uh fireside chat. Um we're going to I have with me uh Theik Shihipa and Cat Woo from Anthropic. We are going to be diving deep into claude code and we'll probably talk a little about this Fable thing that's been going on that that's that's been out there in the news. Actually, on the subject of Fable, literally a minute and a half ago, Fable came back. Like Fable is now available to me. So, if you all want to run out of the room and start using up your Fable credits, I I wouldn't hold that against you. But we're going to have a great conversation. So, please please stick around. Um but yeah, so um please welcome uh The Cat for me. Thanks for having us. Yeah, we timed it for the the chat for sure. Yeah.

欢迎来到这场炉边对谈。今天和我坐在一起的是来自 Anthropic 的 Thariq Shihipar 和 Cat Wu。我们会深入聊聊 Claude Code,可能也会聊聊最近新闻里沸沸扬扬的 Fable。说到 Fable——就在一分半钟之前,Fable 恢复可用了,我现在能用了。所以如果各位想冲出会场去消耗你们的 Fable 额度,我完全不会怪你们。不过接下来会是一场很棒的对话,请大家务必留下来。好,欢迎 Thariq 和 Cat。(Cat/Thariq:谢谢邀请。是啊,我们特意把发布时间掐在这场对谈上。)


[0:57] Simon

Yep. Yep. This is this is why it's all happening. Um this year has been somewhat absurd. Uh I it's amazing. Claude Code came out in February of last year. It's under a year and a half old and it was a bullet point on the Claude Sonnet 3.7 launch. I'd love to hear from you. How has your how has what you do on a day-to-day basis changed in the past year now that we have these coding agents that actually work for us?

对对,这就是一切发生的原因。这一年简直有点荒诞。太惊人了——Claude Code 是去年二月发布的,到现在还不到一年半,而且当初它只是 Claude Sonnet 3.7 发布公告里的一个要点(bullet point)。我很想听听:现在有了这些真正能替我们干活的 coding agent,过去一年你们的日常工作发生了什么变化?


[1:22] Cat

I remember when we first came out with Cloud Code and Sonnet 37, you would give it this task and you would have to closely monitor every single little thing that it tried to do. Uh I remember I would read every permission prompt extremely carefully. I would frequently say no. I would always say no no no like did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like we've all gotten a chance to just take a step back, delegate a lot more of the like menial implementation to claude. And it just freed up a lot of our time to think about more creative work like what is the right experience that we should be providing to our users now that we know cloud code can implement a lot of it. And now with Fable, it's just a totally different step change improvement. um we see for a lot of our use cases that you can actually oneshot a ton of features with Fable now. So it it's been amazing to see the transition and to go through this with all of you in the community.

我记得我们刚推出 Claude Code 和 Sonnet 3.7 的时候,你给它一个任务,然后必须紧盯它做的每一件小事。我记得那时候我会极其仔细地读每一条权限确认提示,经常直接拒绝,总是在说'不行不行不行'——你检查这个文件了吗?检查那个文件了吗?而现在,每一代模型的进步都令人难以置信。我感觉我们所有人都得以退后一步,把大量琐碎的实现工作交给 Claude,这就腾出了很多时间去思考更有创造性的问题,比如:既然我们知道 Claude Code 能把大部分东西实现出来,那我们到底应该给用户提供什么样的体验?而到了 Fable,又是一次完全不同量级的跃升——在我们的很多使用场景里,现在用 Fable 真的可以一把梭(one-shot)搞定一大堆功能。能亲历这个转变,并且和社区里的各位一起走过来,实在太棒了。


[2:21] Thariq

Yeah, I mean I I think I remember the first text I got about cloud code. One of my best friends was like, "Oh, you need to go try Claude Code." And it was about just when Opus 4 came out and I I tried it and I was like, "Oh, shit." Like I I need to work at Anthropic now, you know? And that was Opusport which I mean great model but like yeah you were permission prompts and uh yeah I think it's kind of crazy how much amnesia we have I think where I'm like oh like auto mode has always been here right like I I don't even remember pressing like yes and allow. Um and yeah I think for me the big thing that I'm trying to push myself is like oh we have to do like high higher quality work than we've ever done before. You know like the the outputs are like like incredibly high quality. like I've been using it to edit videos a bunch and I'm like okay it has to meet the very exacting demands of our brand team and in a couple hours or we just can't do it you know and so um yeah I think that's how I'm sort of like trying to shift with Fable where it's like okay the best work we've ever done like faster than we've ever done it before

是啊,我还记得我收到的第一条关于 Claude Code 的短信。我一个最好的朋友说:'你必须去试试 Claude Code。'那大概正是 Opus 4 发布的时候,我试了之后心想:'我去,我得去 Anthropic 工作了。'那还是 Opus 4——当然是很棒的模型,但那时候你还得应付各种权限确认。我觉得很有意思的是我们的'失忆症'有多严重——现在我会觉得'auto 模式不是一直都有吗',我甚至都不记得自己按过'允许'了。对我来说,现在最想推动自己的一点是:我们必须做出比以往任何时候都更高质量的工作。因为产出的质量已经高到离谱。比如我最近一直在用它剪视频,我的标准是:它必须在几个小时内达到我们品牌团队极其苛刻的要求,否则这事就干脆做不成。所以这就是我在 Fable 时代给自己的定位:做我们做过的最好的工作,而且比以往任何时候都快。


[3:18] Simon

I've certainly been finding that myself like software engineering is getting harder because the the level of ambition of the th stuff we can take on has gone up like I I have such higher expectations of myself now that I I have these tools to back me up which is fun but it's al it's a lot of work.

我自己也确实有这种感受:软件工程反而变难了,因为我们能挑战的事情的野心水位涨上去了。有了这些工具撑腰,我对自己的期望高了太多——这很有趣,但也真的很累。


[3:34]

Yeah. It's all the thinking is you know it's it's tiring.

是啊,全是脑力活……确实挺累的。


[3:38] Simon / Cat

And so what's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world? I think one of the biggest shifts that we're seeing in the ENG skill set is I think you know two years ago it was pretty typical for a product manager to go talk to a bunch of customers and over the course of six months align with like cross functional teams on some PRD and then write this like thorough spec um and uh dock on how exactly we'll implement this before the first line of code gets written and now things are like completely turned the opposite way. I I think for a lot of engineers, the the push I would give to a lot of folks in the room is to develop more of your business sense and product sense on what is it that we should build because now that this the timeline between having this idea and building it is so much shorter. It's down from six to 12 months to maybe even a week. That means all of us need to have better taste on what is it that is worth building. what is it that will actually inflect the businesses that we're working on. So I think it's like an increase in value on product taste and business sense and a bit lower on execution in in most in most product domains. Of course for infra uh there there's still a very heavy emphasis on making sure all the details are right.

(Simon)那么,有哪条一年前还成立的软件工程'常识',你们觉得在这个新世界里已经不再成立了?(Cat)我认为工程师技能组合里最大的一个转变是:两年前的典型流程是,产品经理去访谈一堆客户,花六个月和各个跨职能团队对齐出一份 PRD,然后写出详尽的规格说明和实现文档,之后才写下第一行代码。而现在事情完全反过来了。对在座的很多工程师,我想给的建议是:去培养你的商业嗅觉和产品感——到底该做什么。因为从产生想法到把它做出来的周期已经短太多了,从过去的 6 到 12 个月缩短到可能只要一周。这意味着我们所有人都需要有更好的品味:什么东西值得做?什么东西能真正撬动我们所服务的业务?所以我认为产品品味和商业嗅觉的价值在上升,而在大多数产品领域,纯执行的价值在相对下降。当然,对 infra 来说,把所有细节做对依然极其重要。


[5:09] Thariq

Yeah. I I think for me uh it's like rewrites are now good. You know what I Like I I think that like

对我来说,变化是:重写(rewrite)现在是件好事了。你懂我意思吧,就是说——


[5:15] Simon

the worst thing you could do is now actually fine.

以前你能干的最糟糕的事,现在居然是没问题的了。


[5:18] Thariq

Yeah. Exactly. Exactly. Like all the like especially Yeah. mythical man stuff like never rewrite like I I'm a pro rewriting now you know like if you have a good test suite I think actually the rewrite forces you to like make sure you have a good test suite but

对,正是如此。尤其是《人月神话》那套'永远不要重写'的教条——我现在是坚定的重写派。当然前提是你得有一套好的测试套件,而且我觉得重写这件事本身会倒逼你把测试套件建好。


[5:31] Thariq

um I think that like what people I think underount is like a codebase is a spec and maybe it's the only copy of the spec that you have right because like no one knows every branching part of the codebase and and yeah you can take this as like an artifact and like distill it or create other versions of it obviously like yeah we rewrote bun in rust And uh you know it works great like you know it's it's live for me right now. Um

我觉得大家低估了一点:代码库本身就是一份规格说明(spec),而且它可能是你手上唯一的一份——因为没有人真正了解代码库里的每一条分支逻辑。你可以把它当作一件'制品',去蒸馏它,或者派生出其他版本。很明显的例子:我们用 Rust 重写了 Bun,效果很好——我现在跑的就是那个版本。


[5:54] Simon

but you're not shipping clawed code on bun in rust yet right or is that

不过你们还没把 Rust 重写版 Bun 上的 Claude Code 正式发布吧?还是说……


[5:58] Thariq

internally we have.

内部已经在用了。


[5:59] Simon

Wow. Oh that's exciting. Yeah.

哇,这太让人兴奋了。


[6:01] Thariq

Yeah. Yeah.

是的,是的。


[6:04] Simon

I think what you're saying there about um yeah the uh I'm sorry I just lost my train of thought. But yeah, the um the rewrites thing has been really interesting for me as well because you can almost come up with a good test suite and then spin up three implementations and pick which of those implementations was the most accurate. I'm doing a lot more prototyping now. I've always been a prototyper and now I pro I prototype things on my phone during the conference just so I've got something that I can pick up later on and that's working now which is kind of extraordinary. So the other big launch recently was claude tag which that's what a week old now I think at least for the rest of us. Claude tag I understand that's being used in anthropic by non by non-engineers a great deal. What kind of things are non-engineers doing with claude tag?

你刚才说的那个……抱歉,我一下子断片了。不过重写这件事对我来说也特别有意思:你几乎可以先搞出一套好的测试,然后同时开三个实现,最后挑出其中最准确的那个。我现在做的原型也比以前多得多——我一直是个爱做原型的人,现在我甚至会在开会期间用手机做原型,就为了留个东西回头可以捡起来继续——而这现在真能跑通,简直不可思议。另一个最近的大发布是 Claude Tag,对我们这些外部用户来说才一周左右吧。我听说 Claude Tag 在 Anthropic 内部被大量非工程师使用。非工程师们都在用 Claude Tag 干什么?


[6:51] Cat

So claw tag is a claw that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about claw tag is it's multiplayer by default. And so once you add cloud tag into a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. So you can tell claw tag, hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase, and it'll do it for the lifetime of the channel without you having to manually tag it in. And then the third big shift that we've seen is um we've added team memory into this. So if you tell quad tag your preferences in the channel, it'll remember this for every future post. So if you wanted to um always debug like if you always wanted to debug outages, but you don't want it to debug warnings, um just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team. Internally we see quad tag as the evolution of quad code. So we see this as a large shift in how we work internally. Um claw tag currently lands 65% of our product PRs

Claude Tag 是一个住在你们团队协作工具里的 Claude。我们上周先在 Slack 里发布了它。Claude Tag 的不同之处在于,它默认就是多人协作(multiplayer)的:一旦你把 Claude Tag 加进一个 Slack 频道,你可以插话,你的队友也可以插话,大家可以一起协作推进同一个 PR。第二个大区别是它是主动的而不是被动的。你可以告诉 Claude Tag:'嘿,盯着这个频道里的每一条 bug 报告,提一个 PR 修掉它,并且 @ 最近改过这块代码的工程师'——它会在这个频道的整个生命周期里持续这么做,不需要你每次手动 @ 它。第三个大变化是我们加入了团队记忆(team memory):如果你在频道里告诉 Claude Tag 你的偏好,它会记住并应用到之后的每一条消息。比如你希望它总是去排查线上故障,但不要去处理 warning,用自然语言在频道里说一句就行,它会替你和团队里的每个人都记住。在内部,我们把 Claude Tag 视为 Claude Code 的进化形态,这是我们内部工作方式的一次重大转变。Claude Tag 目前承包了我们 65% 的产品 PR。


[8:13] Simon

for for all of anthropic and for cloud code or just for cloud code.

是整个 Anthropic 的,还是只算 Claude Code 团队的?


[8:18] Cat

Uh this is just for our product engineering team. So our internal version of claw tag lands 65% of our product PRs right now. And this is a huge shift. This is like this is more than 50% of our PRs. Um, and the way that we actually see people split work between cloud code and cloud tag is cloud code is still the best place for your most complex tasks when you're interactively iterating with the agent. But claw tag is great for having it work proactively on your behalf so that you no longer need to um manually uh kick off quad codes for for all of the bug reports that might come up for features that you're working on.

呃,这是我们产品工程团队的数据。也就是说,我们内部版的 Claude Tag 目前落地了我们 65% 的产品 PR——这是个巨大的转变,超过了一半。至于大家在 Claude Code 和 Claude Tag 之间如何分工:Claude Code 仍然是处理最复杂任务的最佳场所,适合你和 agent 交互式地反复打磨;而 Claude Tag 擅长主动替你干活,这样你就不用再为你负责的功能冒出来的每个 bug 报告手动去开一个 Claude Code 会话了。


[9:00] Thariq

Yeah. And for like non-coding cases like I I I think we've seen people use claw tag just uh like for example before this before this talk we asked claw tag like hey when is Fable releasing? We wanted to make sure that like you know we'd line it up with the announcement. Um, and so Cloud Tag would search our Slack and and look at, you know, who who's been saying what. Uh, so as a search engine for your company is really valuable. Uh, it has all the context for your product. So you can ask it like metrics related questions. And oftent times when you're making decisions, you want it to be informed by like, you know, what do the metrics say? And then so like you hook it up to your event store. Um, I've seen like our marketing team do things like, oh, like, hey, tell me about this feature. And, you know, they're not programmers, but Claude is a programmer. I can clone the code base and be like, "Oh yeah, this is like, you know, the feature. This is what it looks like. This is a recording of me using the feature, you know. So like um yeah, it just enables a whole wide variety of things." And I think we're still early on in figuring that out. Yeah.

对,还有非编程的场景。比如这场对谈开始前,我们就问了 Claude Tag:'Fable 什么时候发布?'——我们想确保对谈和发布时间对得上。Claude Tag 会去搜我们的 Slack,看谁说过什么。所以把它当作公司内部的搜索引擎非常有价值。它掌握你们产品的全部上下文,你可以问它指标相关的问题——毕竟做决策的时候你往往希望有数据支撑,那就把它接到你们的事件存储上。我还见过我们市场团队的人问它:'嘿,给我讲讲这个功能。'他们不是程序员,但 Claude 是啊——它可以把代码库 clone 下来,然后说:'这就是那个功能,长这样,这是我使用这个功能的录屏。'它就是能解锁各种各样的用法,而且我觉得我们还处在探索这些用法的早期。


[9:54] Simon

Well, I feel like this is one of the fascinating things about the Claude Code story is you use Claude Code to build Claude Code and you've been doing this since presumably before the public launch of launch of Claude Code a year and a half ago. And yeah, one of the problems I've had with coding agents is I get how to use them as an individual, but I'm not really clear on how I use that in a team in a team environment. It sounds like Claude Tag is is your current answer to that sort of team collaborative layer for this stuff.

我觉得这正是 Claude Code 故事里最迷人的地方之一:你们用 Claude Code 来构建 Claude Code,而且大概从一年半前公开发布之前就开始这么干了。我用 coding agent 遇到的一个问题是:作为个人我知道怎么用,但在团队环境里该怎么用,我一直没想明白。听起来 Claude Tag 就是你们目前对'团队协作层'这个问题的答案。


[10:18] Cat / Thariq

Exactly. And a large percentage of our sessions are actually multiplayer right now. So that means maybe I say, "Hey, I think we should like um implement this new feature in co-work." And I'll tag in claw tag to do a first pass at it. And then I'll tell cloud tag, hey, just like share a recording of of your final implementation. And then I'll tag in design to take a look and they'll nudge it and then they'll pass it on to edge to like take it to the finish line and get it out to prod. And so it's been this like very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session, but we found that people just like observe how others use it and then follow those social norms and it's actually been pretty easy for pretty intuitive for us to integrate claw tag into our teams. Yeah, I I think it's great for like Yeah. teaching people and also like kind of kind of reducing slo because because you know like if you see someone just be like hey at Claude fix this or something you're like uh you know like I I think there's like some societal or not societal like just like uh like the fact that everyone is seeing you use cloud together sort of levels up how you use claude as well. So

(Cat)正是如此。而且我们现在有相当大比例的会话实际上是多人协作的。比如我说:'嘿,我觉得我们应该在 Cowork 里实现这个新功能。'然后我把 Claude Tag 拉进来先做第一版,再告诉它:'把你最终实现的录屏分享出来。'接着我 @ 设计师来看一眼,他们微调一下,再交给工程把它推到生产环境收尾。整个过程非常流畅。我们还在摸索'多人共同引导同一个会话'的社交规则,但我们发现大家会观察别人怎么用,然后自然地跟随这些社交规范——所以把 Claude Tag 融入团队其实相当顺畅、相当直觉。(Thariq)对,我觉得它还特别适合'教会'大家,也能减少偷懒糊弄——因为如果别人看到你只是丢一句'@Claude 修一下这个',大家心里是有数的。所有人都能看到彼此怎么用 Claude,这件事本身就会把整个团队使用 Claude 的水平往上带。


[11:30] Simon / Cat

right you you want to do work that you're proud to do in public that where the quality doesn't doesn't fall off the cliff. So since we're talking about using claude to build CL well how the process for building claude how do you deal with the hardest problem in all of engineering it's prioritization right how do you decide which features are worth building and shipping when building a feature is so much more inexpensive now this is the hard thing so there's a few ways we approach it one is we dog food our products every single day whenever there's something that we want to be able to do in our products that we're not able to instead of finding a different solution we fix our product uh so that it can support this case. Uh we have a very heavy dog fooding culture internally. So before we are able to share our products with everyone in the world, we share it with everyone within anthropic and we share it with some early customers who give us very honest feedback about it. The more brutal the better and we iterate until people love it. So we we have a internal bar for the number of active users and the the amount of retention a feature has to have before we share it with the world. And because this bar is very clear, every engineer knows what they're trying to hit. And I think this also levels up our polish because if the feature isn't polished, people will churn and then we shouldn't ship that feature.

(Simon)对,你会想在公开场合做拿得出手的工作,质量不能垮。既然聊到用 Claude 来构建 Claude——那你们怎么解决工程界最难的问题:优先级排序?当做一个功能变得这么便宜的时候,你们怎么决定哪些功能值得做、值得发布?(Cat)这确实是难点。我们有几种做法。第一,我们每天都在'吃自己的狗粮'(dogfooding):只要我们想在自己产品里做到某件事却做不到,我们不会去找别的方案,而是把自己的产品修好,让它能支持这个场景。我们内部有非常浓厚的 dogfooding 文化。在把产品分享给全世界之前,我们会先分享给 Anthropic 的每一个人,以及一些愿意给出非常坦率反馈的早期客户——越毒舌越好——然后一直迭代到大家真心喜欢为止。我们内部有一条明确的门槛:一个功能必须达到一定的活跃用户数和留存率才能对外发布。因为这条线足够清晰,每个工程师都知道自己要冲的目标是什么。我觉得这也拉高了我们的打磨水准——因为功能不够精致,用户就会流失,那这个功能就不该发。


[12:54] Simon

Do you have an example of a feature which surprised you? you you rolled it out and the engagement was off the charts and it it became was something that was unlikely to be shipped that actually turned into a a real product thing.

有没有哪个功能让你们特别意外?就是那种推出之后使用量爆表、本来根本不太可能被发布、结果却变成正经产品功能的例子。


[13:06] Cat / Simon

I do have one. So a lot of folks on our team love remote control. So remote control lets you connect use your mobile device or quad in the web browser to connect a local quad code session running in your CLI. I never have this need because I just kick off the task directly on mobile and it runs in a cloud session and doesn't use my local environment. I think this is because I'm doing very easy coding tasks. But this is something where I didn't totally understand it. I was like, hey, people should just set up their remote dev environments. But in practice, once we rolled out remote control, everyone like so many people who I talk to are like, "Okay, now what I do every night is I um a lot of people tell me that they just plug their lo their laptop into a power charger, close the screen or like open a bunch of remote control sessions, lock the screen, and then use their mobile phone from their couch to to control quad code." And so this has been this flow that we're now leaning into that I didn't originally get, but now I do. I do exactly that. I get so much so much work on my laptop done from more comfortable environments because yeah, I can remote control it now. That's really fun. Um, how does code review work? Are you pres are you reviewing does a human being review every line of production code that makes it into clawed code? And if not, what are you doing? How do you keep the quality up?

(Cat)还真有一个。我们团队很多人特别喜欢 remote control(远程控制)功能:它让你用手机或者网页版 Claude 连上一个跑在你本地 CLI 里的 Claude Code 会话。我自己从来没有这个需求,因为我直接在手机上发起任务,它跑在云端会话里,不占用我的本地环境——大概因为我做的都是很简单的编码任务。所以一开始我完全不理解这个功能,我想:大家去配远程开发环境不就好了?但实际推出 remote control 之后,我聊到的很多人都说:'我现在每天晚上的固定动作是,把笔记本插上电源,开好一堆 remote control 会话,合上盖子锁屏,然后瘫在沙发上用手机遥控 Claude Code。'于是这个我一开始没 get 到的用法,现在成了我们主动押注的工作流——我自己现在也完全是这么干的。(Simon)我用笔记本干的活里有太多是在更舒服的环境里完成的,因为现在能远程控制它了,真的很爽。那 code review 是怎么运作的?进入 Claude Code 生产环境的每一行代码都有人类审过吗?如果没有,你们靠什么保证质量?


[14:32] Thariq

Sure. Yeah. Um, it it varies on the task a lot. So, so we um for important areas we have code owners, right? And so, uh the system prompt is kind of an example where we have a code owner uh you really need to like uh submit your uh you know you need to get their approval. Um and then uh

好的。这个很大程度上取决于任务本身。对于重要的区域我们设有 code owner(代码负责人)。比如系统提示词(system prompt)就是一个例子——它有专门的 code owner,你必须拿到他们的批准才行。


[14:49] Simon

so I guess the code anyway is directly responsible for the quality of that area of the code.

所以这个 code owner 就直接对那块代码的质量负责。


[14:53] Thariq

That's right. Yeah. Yeah.

没错,是的。


[14:54] Simon

And they need to approve the PR that touches it.

任何动到那块代码的 PR 都需要他们批准。


[14:56] Thariq

Right.

对。


[14:57] Thariq

That's right. uh we have code review our like uh our uh code review git b you know review everything and and so um that like goes on every PR and often times like that's doing the bulk of the review. Um I think that like uh we I something I've seen on the team is like for more complex PRs you might make like an artifact to explain the PR so that other people can then review. Um and uh yeah we just invest a lot into verification CI/CD things like that to make sure that like you know uh anytime anything fails like we have a test we have like uh a really robust environment cloud can control cloud code and test it you know what I mean so um yeah there is a uh yeah there's just like a multi-pronged approach to like code review I think you have anything to add

对。我们有 code review——我们的代码审查机器人,会把所有东西都过一遍,每个 PR 都跑,很多时候审查的大头就是它干的。我在团队里还看到一个做法:碰到比较复杂的 PR,作者会做一个 artifact 来解释这个 PR,方便其他人来审。另外我们在验证、CI/CD 这些方面投入很大,确保任何东西一挂就有对应的测试兜住。我们还有一套非常健壮的环境,Claude 可以操控 Claude Code 来测它自己,你懂我意思吧。所以说,code review 是多管齐下的。Cat 你有要补充的吗?


[15:46] Cat

in general we are trying to move to a world where humans don't need to be in the loop And so for the most critical core changes to the core of quad code and other the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly for the changes that are at the outer layers, we actually have quad code review fully review those. That sounds pretty scary but there we've had this six plus monthl long process to get here and I think there are like baby steps that you take to build up trust with code review. So in the beginning we would have human review for everything and then increasingly we would say okay for code changes that touch these files code review is catching a 100% of the issues there. So we actually don't need a human to be manually reviewing those. Um and then also when we have incident review, we look at the PRs that cause the incident and we say okay how do we update code review to catch that and then we also take those PRs and add it to an eval set to make sure that our future changes to code review never regress that metric. So it it is a big like removing humans from the code review loop is a big step forward. I think it can sound scary and it's not something that you can do overnight, but it is something that you can do through like many months of investment in the infrastructure to give you the confidence that code review is catching everything that you care about.

总体上,我们在努力走向一个人类不需要在环路里(in the loop)的世界。对于最核心、最关键的改动——比如 Claude Code 的核心和其他产品的核心——始终有 code owner,他们会人工审查所有变更。但对于外围层的改动,我们现在越来越多地让 Claude 的 code review 完全接管审查。这听起来挺吓人的,但我们花了六个多月才走到这一步,这中间有一系列小步快跑来建立对 code review 的信任。一开始所有东西都是人工审;后来我们逐步确认:好,对于只碰这些文件的代码改动,code review 能抓住 100% 的问题,那这部分就不需要人再手动审了。另外每次做事故复盘(incident review)时,我们会回看引发事故的那些 PR,思考怎么更新 code review 才能抓住这类问题,然后把这些 PR 加进 eval 集,确保将来对 code review 的任何改动都不会在这个指标上回退。把人从代码审查环路里拿掉是很大的一步,听起来可怕,也不可能一夜做到,但通过好几个月在基础设施上的投入,你可以获得足够的信心:你在乎的东西 code review 全都能抓住。


[17:16] Simon

So, it's interesting you mentioned building trust in the models because that's something I found is like I know that Opus 4.8 if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, it's just going to get it right. Like that that that's not something I have to review closely. But then a new model comes along and I still don't know how do I build trust in fable quickly that it's not going to mess things up that Opus didn't. Is that something that you you you have to think about much the those new model like how does the new model affect your intuition for for what it can do and what it can't do? So the main reason that we're building up this eval base over time is so that new models can be a drop-in replacement because what we do when we have a new model is we run the whole eval set and we make sure that for example Fable is strictly better than Opus 48 and that gives us the confidence to drop it in.

有意思,你提到对模型建立信任——这正是我自己的体会。我知道让 Opus 4.8 给我写一个跑 SQL 查询再输出 JSON 的 endpoint,它肯定能做对,这种活我不用细审。但新模型一出来,我又不知道该怎么快速建立对 Fable 的信任——凭什么相信它不会把 Opus 做得好的地方搞砸?这是你们要重点考虑的事吗?新模型出来时,它会怎么影响你们对它能干什么、不能干什么的直觉?(Cat:)我们长期积累这套 eval 基础,主要原因就是为了让新模型可以直接无缝替换(drop-in replacement)。新模型来了以后,我们会跑完整个 eval 集,确保比如 Fable 严格优于 Opus 4.8,这就给了我们直接换上去的信心。


[18:04] Simon

And are those model evals for Anthropic as a whole or are these claud code team specific evals that you're using?

这些模型 eval 是 Anthropic 全公司层面的,还是 Claude Code 团队自己专用的?


[18:11] Cat

We have both. So we we have evals on our team and we run code review across every repo within Anthropic and so we have evals for that and for things like auto mode we not only have evals across every user within Enthropic but we've also commissioned multiple external testers to red team this to create environments with prompt injections and malicious inputs and make sure that automode doesn't let any of those pass. So for claude code itself and this is a challenge I've had with stuff I'm building. I want to know if the system prompt improvement I made actually improved the product. Right? That's that's the sort of most basic form of product specific eval. And I still don't have a great feel for how to do that. Is something is that something that you're doing such that you have complete confidence that this tweak that you've made to the system prompt does result in in better in better output? We don't have complete confidence but we do a lot to make sure that we don't regress performance. So the starting point that we have is we have a suite of external evals that we trust and we complement that with an even larger suite of internal evals that we trust. Uh to start we mainly optimize for capability. So given given a complete definition of a task and the full codebase does claude make the right decisions and fully fix the the bugs and pass all the tests. So that's the starting point and that's the thing that we optimize for because it is like most directly what users want. But there's a lot of like behaviors that impact how users feel when they work with quad code. For example, people really don't like it when cloud code says it's like time to go to sleep. or people really don't like it when it says like, "Hey, I finished two out of five parts. Like, do you want me to continue?" Like, "Yes, please continue." Um, and so we're building up a set of behavioral emails to catch these. And as we get user feedback, please be loud with us about your user feedback. As we get user feedback, we just rank, okay, these are the priority issues. And we go down one by one and build evals for each of them. So, it's not 100% coverage, but we try to it it is a priority for us to increase the coverage. And how much overlap is there between the how much interaction is there between the Claude code team and the teams at Anthropic who are training the models in the first place? Is is that quite a close collaboration? Now

两种都有。我们团队自己有 eval,而且我们的 code review 跑在 Anthropic 内部所有仓库上,所以有针对它的 eval。像 auto mode 这种功能,我们不仅在 Anthropic 内部所有用户身上做了 eval,还请了多家外部测试方来做红队测试(red team)——构造带 prompt injection 和恶意输入的环境,确保 auto mode 一个都不放过。(Simon:)那 Claude Code 本身呢?这是我自己做东西时一直头疼的:我想知道我对 system prompt 的一处改动到底有没有让产品变好——这算是最基础的产品级 eval 了,但我到现在也没摸出好办法。你们做到了吗?改一处 system prompt,就能完全确信输出确实变好了?(Cat:)不敢说完全确信,但我们做了大量工作来保证性能不回退。基础是一套我们信任的外部 eval,再加上一套规模更大的、我们信任的内部 eval。一开始我们主要优化能力(capability):给定完整的任务定义和整个代码库,Claude 能不能做出正确决策、彻底修好 bug、通过所有测试。这是起点,也是我们优化的重心,因为它最直接对应用户想要的东西。但还有很多行为会影响用户和 Claude Code 协作时的感受。比如大家特别讨厌 Claude Code 说"该睡觉了",也特别讨厌它说"我完成了五步里的两步,要我继续吗?"——废话,请继续啊。所以我们在搭一套行为类 eval 来抓这些问题。收到用户反馈后——请大家大声把反馈砸过来——我们会排优先级:这些是最重要的问题,然后一个一个往下做,给每个问题建 eval。覆盖率还没到 100%,但提高覆盖率是我们的优先事项。(Simon:)那 Claude Code 团队和 Anthropic 负责训练模型的团队之间,交集有多大?合作算紧密吗?


[20:30] Cat

across Anthropic, we all work quite closely together. So um we we meet often to talk about like what do we expect the next generation of models to be able to do. I think our research team has also been amazing about just like showing this publicly. So we often talk in our blog posts about how we're targeting everinccreasing longer horizon work, how we train quad itself to be honest, harmless, and helpful. Um, we also put a lot of effort into making sure that it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want. So Quad um has all the context, but even when you're you're not specific, uh we we teach Claude to make good assumptions and yeah, I I think it's been a it's been a productive partnership.

在 Anthropic 内部,大家合作都非常紧密。我们经常开会讨论:我们预期下一代模型能做到什么。我们的研究团队也很棒,把很多东西公开出来了——我们的博客里常写到我们在瞄准越来越长时程(longer horizon)的工作,写到我们怎么把 Claude 训练得诚实、无害、有帮助。我们也花了很大力气确保它对齐你的意图,哪怕你的意图表达得很模糊。当然,你最好还是尽量把要求说具体,让 Claude 拿到完整上下文;但即使你说得不具体,我们也教 Claude 去做合理的假设。总之这是一段很高产的合作关系。


[21:22] Simon

And so um Derek, you this morning you mentioned that the system prompt for Claude Code has been reduced by 80%. Because of Claude Fable, can you go into a little bit more detail about what that looks like? What kind of things have you been able to drop? Yeah. So it wasn't just Fable, it was uh Opus 4.8 as well. And um yeah, going forward the future models, but we do sort of um we have different system prompts for different models now. Um I think that like uh some of the patterns we saw is that um we were over constraining Claude, right? So I think the initial like maybe Opus 4ish kind of models wanted a lot of examples and uh removing examples was extremely helpful because it was just more creative than like uh you know the examples we gave it.

Thariq,你今天上午提到 Claude Code 的 system prompt 因为 Claude Fable 砍掉了 80%。能不能展开讲讲这具体是什么样?哪些东西可以直接删掉了?(Thariq:)其实不只是 Fable,Opus 4.8 也一样,未来的模型也会如此。我们现在是不同模型用不同的 system prompt。我们看到的一个模式是:我们以前把 Claude 约束得太死了。早期大概 Opus 4 那一代的模型需要大量示例(examples),而现在把示例删掉反而特别有帮助——因为模型自己比我们给的示例更有创造力。


[22:06] Simon

That's really because one of the top prompting tips I give people is give it examples like examples are the easiest ways. If that's no longer true that kind of breaks my prompting model a little bit.

这太有意思了,因为我给别人的头号提示词建议就是"多给示例"——示例是最省事的办法。如果这条不再成立,那我的提示词心智模型就有点被打破了。


[22:16] Thariq

Ex yeah same same here. I think I was surprised to hear that. I I think that now it's more about like sort of the shape of what you give the tools to Claude and like yeah your system prompt and things like that. Um the other thing we did is we um we we try and give it more context and fewer like do not do this you know because like I I think that um it's just a very strong impulse to Claude and especially if that uh conflicts with user instructions later on that can be like extremely confusing to Claude, right? because you're like, "Oh, like I've got this skill that says this and the system prompt says this." Um, and so we try and like uh have fewer hard constraints and more just like sort of context and just like fewer instructions overall. Um, yeah, I think it's definitely a science. It took a bunch of like evals to build. I'm not sure if you had anything else on uh the lean system.

我也一样,听到这个结论时我也很意外。我觉得现在更重要的是你给 Claude 的工具的"形状",还有 system prompt 的整体结构这类东西。另一个改动是:我们尽量多给上下文,少写"不许做这个"之类的禁令。因为这种禁令对 Claude 来说是很强的冲动,一旦后面和用户指令冲突,Claude 会非常困惑——"我这个 skill 说要这样,system prompt 又说要那样"。所以我们尽量减少硬性约束,多给上下文,整体指令也更少。这绝对是一门科学,靠一堆 eval 才打磨出来的。Cat,精简 system prompt 这块你还有要补充的吗?


[23:06] Cat

I think in general when you're prompting these models, you should always think about like are there edge cases to the instruction that I'm giving it? And when we went back and we reviewed all the instructions in the cloud code system prompt, we found a few cases where yes, this statement is like 90% true, but there's like a real 10% of cases where this is not true. And we didn't want to constrain the model or like confuse it into thinking, hey, it should always do this. Like one good example is verification. Everyone here wants claw to verify its work. Um, and we had some instructions in the prompt that just said um, if you make a front-end change, always verify. But you know, there there is a limit to it. Like for example, if you if it's changing copy from one string to another string and the user tells says it says like just make a quick fix and update the test, maybe you don't want to verify. And so we we've also adjusted our wording from saying always verify verify verify verify to like hey most of the time when you're doing front-end work you can't always understand the full experience by hitting the backend endpoints. So like when when you make like more changes to the user experience uh please run the app locally and actually in fact that instruction probably isn't even good because what is a large change like maybe it wants to change it test it for small changes too. Um, in general, whenever you give a prompt to the model, you should always think about the ways in which it could be misinterpreted by like a well-intentioned uh other user or human in order to better understand how the model might interpret it and in order to make sure that you can soften the prompt such that it is actually 100% accurate because you are giving this prompt to the model 100% of the time. But what's fascinating about that is you're relying on the model's judgment. And that's got to be a opus fable level thing. Like models a year ago did not have the levels of judgment necessary to decide if they were going to like test a change or not. That's absolutely fascinating. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks and stuff.

总的来说,给这些模型写提示词时,你要时刻想:我给的这条指令有没有边界情况(edge case)?我们回头把 Claude Code system prompt 里所有指令都过了一遍,发现好几处是"这话 90% 的情况成立,但确实有 10% 的情况不成立",我们不想因此把模型框死、让它误以为永远都该这么做。一个很好的例子是验证(verification)。在座每个人都希望 Claude 验证自己的工作。我们的 prompt 里原来写着:只要改了前端,一律去验证。但这是有限度的——比如只是把一个文案字符串改成另一个,用户说"就是个小修复,顺便更新一下测试",那可能就不需要验证。所以我们把措辞从"永远要验证、验证、验证"调成了"前端工作大多数时候光打后端 endpoint 是看不全用户体验的,所以当你做的改动对用户体验影响较大时,请在本地跑起应用来实际看一看"。其实这条指令可能都还不够好——什么叫"较大的改动"?也许小改动它也应该测。总之,每次给模型写 prompt,都该想想它会被一个善意的"另一个人"怎么误读,以此理解模型可能怎么理解它,然后把措辞软化到真正 100% 准确——因为这段 prompt 是 100% 的时间都喂给模型的。(Simon:)这里面最迷人的一点是:你们在依赖模型的判断力。这一定是 Opus、Fable 这个级别才有的东西——一年前的模型根本没有这种判断力去决定要不要测一个改动。太有意思了。不过如果你面向的是一大堆模型、想让便宜任务跑便宜模型,这套打法就不成立了。


[25:16] Cat

We actually have a different system per model now because of this very reason. So, it's only our uh most frontier models that have this 80% token decrease and the older models actually still have the full system prompt.

正因为这个原因,我们现在每个模型用不同的 system prompt。只有最前沿的模型才享受这 80% 的 token 削减,老一些的模型用的还是完整版 system prompt。


[25:28] Simon

Do you think Fable and Opus are smart enough to be able to prompt Haiku with more details because they understand that Haiku has less judgment, has less taste?

你们觉得 Fable 和 Opus 够不够聪明,能不能因为理解 Haiku 判断力更弱、品味更差,就主动给 Haiku 写更详细的提示词?


[25:41] Thariq

We haven't been able to evalu, but we don't have any hard data to show it. I think there's a tough thing with um smaller models sometimes because like uh you know we saw this with um just like sometimes the the larger models can be more token efficient on a hard problem than the smaller models. And so uh you know there's like a little bit of that uh intuition to build about like you know sometimes you really just want frontier intelligence almost all the time you know um but it's uh yeah like the paralle curve shifts you know and so it's hard to find. Yeah.

我们还没能系统评估过,没有硬数据能证明。小模型有个微妙的地方:我们观察到,在难题上大模型有时反而比小模型更省 token。所以这里有个直觉要慢慢建立——很多时候你几乎任何时候都想直接要前沿智能。当然,帕累托曲线在移动,这个平衡点不好找。


[26:17] Simon

I mean that's something I found fascinating. I feel like a year ago I did not trust a model to write a prompt. Like today the good models are very good at prompting. Like a lot of my prompts are written by models which feels absurd but it actually works really well. And something that helped me come to terms with that was thinking about sub agents which is entirely about a clawed model setting up a prompt for another claw model so that it knows what to go and do. Yeah, I think workflows are actually a really good example of this because it's like cloud not just prompting a single sub agent, but it's like pro prompting like the orchestration of many sub aents and each one of them gets like, you know, a very detailed prompt. So, it's like almost like a level above like, you know, just spawning a sub agent. Um, so yeah, it's quite good at that. Yeah, I've also been using on my personal machine like giving it the Gemini API and being like, oh, like here, generate images. And it's so good at it's way less lazy than I am at prompting an image model, you know. So, um yeah, it's just Claude prompting Claude all the way down. Yeah,

这是我觉得特别迷人的一点。一年前我根本不信任模型写提示词;今天好的模型已经非常会写 prompt 了——我很多 prompt 都是模型写的,听着荒谬,但真的很好用。帮我接受这件事的是 subagent 的思路:它本质上就是一个 Claude 给另一个 Claude 写 prompt,告诉它该去干什么。(Thariq:)对,workflow 其实是个更好的例子——Claude 不只是给单个 subagent 写 prompt,而是在编排一整批 subagent,每一个都拿到非常详细的 prompt,相当于比"生成一个 subagent"再高一个层次。它在这方面相当强。我自己在个人机器上还给它接了 Gemini API,让它去生成图片——它给图像模型写 prompt 比我勤快多了。所以说,一路往下全是 Claude 在给 Claude 写 prompt。


[27:14] Cat

I think Claude also wrote the prompt for the workflow tool.

workflow 工具的 prompt 我记得也是 Claude 写的。


[27:18] Simon

For the workflow tool, I've read that prompt. It's a good prompt. I mean, that's actually um a frustration I have with Anthropic generally is you publish the prompts for Claude. There's there's a web page with them on, but you don't include the tool prompts and the Claude code prompts. I still have to run a proxy to intercept them. I would love it if the Claude code prompts were deliberately published because they're the documentation. They're how you know what the tool can do and how it works.

workflow 工具那个 prompt 我读过,写得不错。说到这个,我对 Anthropic 一直有个不满:你们公开了 Claude 的系统提示词,有个网页放着,但不包括工具的 prompt 和 Claude Code 的 prompt——我到现在还得跑代理去截获它们。我特别希望 Claude Code 的 prompt 能被正式公开,因为它们就是文档:你想知道这个工具能干什么、怎么工作,看 prompt 就知道了。


[27:42] Cat

I'll write down that feature request.

我把这条功能需求记下来。


[27:44] Simon

Please do.

拜托一定。


[27:45]

I'll have Claude tag do it.

我让 Claude Tag 去办。


[27:47] Simon

And also the diffs like every now and then I'll do I'll diff the older and the newer prompt and that's how I learn the capabilities of the new model. I'm really looking forward to seeing what this 80% reduction actually looks like.

还有 diff——我时不时会把新旧 prompt 做个对比,这是我学习新模型能力的方式。我真的很期待看到这 80% 的削减到底长什么样。


[27:59] Thariq

Yeah, this is on me. I have to make a post about this in detail. Yeah.

这事儿算我头上,我得写一篇详细讲这个的文章。


[28:03] Simon

So, what's your bar? Let's talk about tools a little bit. Claude code is basically a just bag of to a big bag of tools. What's your bar for introducing a new tool? How do you decide when it's worth doing that additional engineering at that level?

那我们聊聊工具吧。Claude Code 本质上就是一大包工具。你们引入一个新工具的门槛是什么?怎么判断值不值得在这个层面做额外的工程投入?


[28:16] Cat

Do you want to take it? Because you introduced one of the best tools we have.

这题你来答吧?我们最好用的工具之一就是你引入的。


[28:20] Thariq

Yeah. I mean, yeah, it's like my career peaked when I introduced the ask user question tool. I I think so. Um, it it's really hard. I I I think is the the especially for some tools like ask user question is claude's tool to ask you and so it's hard to eval and and sometimes it's more of a user preference thing. So like um especially back then we had fewer evals. It was very ant fooding based um or sorry dog fooding is yeah ant fooding is you know our ant version of that. Um but yeah I mean I think overall we've been trying to trend towards fewer tools. I think the last set of tools we introduced were like the task tool I think um and try and give cloud more general versions to do this. Um

行。可以说我职业生涯的巅峰就是引入了 ask user question 这个工具。这事真的很难。尤其某些工具——比如 ask user question 是 Claude 用来反问你的工具——很难做 eval,很多时候更偏用户偏好。特别是当年我们的 eval 还少,主要靠 ant fooding……哦不对,一般叫 dogfooding(内部试用),ant fooding 是我们蚂蚁版的说法。总体上我们一直在朝"更少的工具"收敛。我们最后引入的一批工具好像是 task 工具,思路是给 Claude 更通用的能力去完成这类事。


[29:02] Simon

right one of the most interesting tools is the file editing tool which but you can have file editing as a tool or you can tell teach it tell it to use said and GP and and do things that way. What's the latest evolution of your file editing tool?

最有意思的工具之一是文件编辑工具。你可以把文件编辑做成一个专门的工具,也可以教模型直接用 sed 和 grep 那一套去改。你们的文件编辑工具最新演化到什么样了?


[29:15] Thariq

Uh I think we still have one but like for example we removed our GP and uh other search tools. glob tools for just like native uh like bash and so uh yeah we still have one I I think this is kind of like I said in my uh talk earlier that the models are kind of like more of a biology than a physics you know and so like you know it's hard to like especially tool design I think is quite hard and I'm not sure if actually if cat disagrees and is like oh like we should no there's like a science to the ebal of it but I'm sort of like yeah tool design is more of an art maybe or like a biology Yeah,

文件编辑工具我们还留着,但比如 grep 和其他搜索工具、glob 工具都删掉了,直接用原生 bash。就像我上午演讲里说的,这些模型更像生物学而不是物理学,工具设计尤其难。我不确定 Cat 会不会反对,说"不对,这里面有 eval 的科学"——我自己的感觉是,工具设计更像一门艺术,或者说生物学。


[29:51] Cat

I think I largely agree, but that I think in general as we introduce more tools, it like in general we try to keep the cardality pretty low and make sure that every tool we add has a distinct function from every other tool so that Claude can very easily distinguish when to call each. for file edit. Actually, the reason that we have file edit is because we can render it because back back in the day we used to show or I guess we still do. Um, so we show people when quad makes a file change and there's this nice like dedicated UI that just says uh do you approve this edit to this file? And the reason that we had a dedicated file edit tool was so that we could deterministically know that quad was making a file so we could show people this nice UI. And for a lot of the new users who are onboarding, I think they still really like this experience. So, we've kept it around. But for a lot of us who are on auto mode right now, or hopefully you're not on YOLO mode. But anyway, for a lot of us right now, I don't think actually it matters and we probably could just remove file edit and we'll be totally fine.

我大体同意。不过总的原则是:随着工具增多,我们尽量把工具数量(cardinality)压得很低,并确保每个新工具的功能和其他所有工具都有明确区分,这样 Claude 才能轻松判断什么时候调用哪个。至于 file edit,我们保留它的真正原因是渲染——从前(其实现在也是),Claude 改文件时我们会展示一个专门设计的 UI,问你"是否批准对这个文件的这处编辑"。有专门的 file edit 工具,我们就能确定性地知道 Claude 在改文件,从而展示这个漂亮的 UI。对很多刚上手的新用户来说,他们真的很喜欢这个体验,所以我们留着它。但对我们这些已经在用 auto mode 的人——希望你们没在用 YOLO mode——其实它可能已经无所谓了,把 file edit 删掉大概也完全没问题。


[30:57] Simon

So, let's talk about auto mode or let's talk about safety and security in general. Like I am deeply aware of the risks of prompt injection and there there are so much bad things can happen if somebody else tells my clawed code what to do. I still mostly run clawed code in yellow mode and feel incredibly guilty about it. What's the advice within anthropic for safely running clawed code? Like what do you tell people to do?

那我们聊聊 auto mode,或者说整体聊聊安全。我对 prompt injection 的风险有深刻认识——一旦别人能指挥我的 Claude Code,能发生的坏事太多了。但我大部分时候还是跑在 YOLO mode 上,并且深感愧疚。Anthropic 内部对安全运行 Claude Code 的建议是什么?你们怎么教大家?


[31:21]

Why not auto mode?

为什么不用 auto mode?


[31:23] Simon

I am starting to use auto mode and I don't understand it enough to get how safe it is. But yeah, that's as of maybe three weeks ago, I'm defaulting to auto mode.

我开始用 auto mode 了,但我对它了解不够,不知道它到底有多安全。不过是的,大概三周前起,我默认用 auto mode 了。


[31:32] Cat

Okay, so broadly within Enthropic, almost every single person uses auto mode. It is the best way to do longunning work in quad code while being safe. We've done extensive bashing. We have thousands of eval. We've commissioned many red team teamers to create adversarial environments in order to trick cloud code into doing bad actions. and we've mitigated every single issue that they found. And so we're going to publish some evals in the in the coming weeks, but we we've pretty much mitigated every attack.

好,总体来说,在 Anthropic 内部几乎每个人都在用 auto mode。它是在 Claude Code 里既安全又能跑长任务的最佳方式。我们做了大量的压力测试,有上千个 eval,还请了很多红队人员专门搭建对抗性环境,想方设法诱骗 Claude Code 去做坏事——他们找出的每一个问题我们都做了缓解。接下来几周我们会发布一些 eval 结果,但基本上可以说,每一种攻击我们都已经防住了。


[32:08] Simon

That is a big claim. That's very exciting. If that holds up,

这话说得可够大的。不过很让人兴奋——如果真能站得住脚的话。


[32:12] Cat

we we will we'll share the evals for it so uh f folks can assess, but we've been extremely diligent about identifying all the ways in which quad might mess up and then updating auto mode to counter it. It doesn't catch 100% of things. Um I don't think any Yeah, that would be way too strong of a claim. But for the main categories of risks that we're concerned about like prompt injection, data exfiltration, um the risks are far lower than the average human reviewer.

我们会把相关的 eval 公开出来,大家可以自己评估。但我们确实非常勤勉地梳理了 Claude 可能出岔子的所有方式,然后不断更新 auto mode 去应对。它不可能拦住 100% 的问题——那样说就太夸张了。但对于我们最关心的几大类风险,比如 prompt injection、数据外泄,它的风险已经远低于一个普通的人类审核者了。


[32:44] Thariq

So, and oh yeah, a little bit on how auto mode works. I I think it's useful to build this mental model. So, whenever Claude is doing a turn, uh there's a or a bash call, uh there's a sonet classifier that is judging the tool and also the context of the conversation, your instruction, right? And so there are some things around like uh permissions which are dependent on your request right so you don't want to give git push like permissions all the time but if you say hey push this to github you want it to do it right and so auto mode will or if you say don't push you want it to deny it right and so auto mode will do that particular thing happens to me a lot where it's like oh like auto mode like I like claw tried to do this because it's very helpful and proactive and auto mode saw like oh like you know uh you know don't do this and it like surfaced it. So it's good at like the dynamic permissions that you yourself give inside of the prompt which I think is really important. Um it also works well with our sandboxing infrastructure because like sandboxing is one of those things where there are so many different edge cases. Uh and it's hard for like us to deterministically follow them. But if you we have a sandbox and something needs to escape the sandbox like a you know a network request auto mode can then look at that request and be like oh hey does this like you know does this make sense right and just allow that in.

对了,稍微讲讲 auto mode 是怎么工作的,我觉得建立这个心智模型挺有用。每当 Claude 执行一个回合、发起一次 bash 调用时,都有一个 Sonnet 分类器在评判这个工具调用,同时也结合整个对话的上下文和你的指令。有些权限是取决于你的请求的——你不会想一直开着 git push 这种权限,但如果你说'嘿,把这个推到 GitHub',你就希望它照做,对吧?反过来,如果你说了'别推',你就希望它拒绝。auto mode 就会这么处理。这种情况我自己就经常遇到:Claude 因为太乐于助人、太主动,想去做某件事,而 auto mode 看到'哦,用户说过别做这个',就把它拦下来提示出来。它很擅长处理你在 prompt 里动态给出的这类权限,我觉得这一点非常重要。它跟我们的沙箱基础设施也配合得很好——沙箱这东西边缘情况太多了,我们很难用确定性规则全部覆盖。但有了沙箱之后,如果有什么东西需要突破沙箱,比如一个网络请求,auto mode 就可以看着这个请求判断:'嗯,这个请求合理吗?'合理就放行。


[34:06] Simon

So I hadn't realized auto mode is interacting with the networking sandbox as well.

我之前还真没意识到 auto mode 也会跟网络沙箱交互。


[34:10] Thariq

Yeah exactly. So it's also part of sandbox. Um yeah

对,没错。它也是沙箱体系的一部分。


[34:13]

it it interacts with any permission prompt that the user would otherwise see.

凡是用户原本会看到的权限确认提示,它都会介入处理。


[34:17] Simon

And how old is auto mode? like I feel like as a feature that I had access to, it's it's only a couple of months old, right?

auto mode 有多久历史了?在我印象里,作为一个我能用上的功能,它也就出来几个月吧?


[34:23] Cat

We've been using it within Anthropic since January.

我们在 Anthropic 内部从一月份就开始用了。


[34:26] Simon

Okay.

哦,这样。


[34:27]

So, we've been hardening it for for quite a while. And it's obviously Anthropic is extremely focused on safety and security. And so, we've been working broadly across our um alignment and safeguards teams in order to enable the rollout internally, build out these eval even more robust before sharing it out with the world. I think my only my main problem with automotive is I don't understand it deeply enough. Like for anything that's looking after my security, I want to know as much as I can about how it works and what it protects me against and what it doesn't so I can decide like how much I can trust it.

(Cat)所以我们已经打磨了相当长的时间。而且显然 Anthropic 极其重视安全,我们和内部的 alignment、safeguards 团队广泛合作,先在内部推开,把这些 eval 做得更扎实,然后才向外界发布。(Simon)我对 auto mode 唯一的、也是最大的问题是:我对它的理解不够深。凡是负责保护我安全的东西,我都想尽可能了解它的工作原理——它能防住什么、防不住什么——这样我才能决定该给它多少信任。


[35:00] Thariq

I I think yeah Dell was working on a post about this. So um a little bit just a little bit more about auto mode. This is also the reason cloud tag is so so good, right? because cloud tag uses auto mode and like you can imagine that like one of like I've heard a lot of like build versus buy questions on on a Slackbot. I'm like please you probably shouldn't build your own AI Slackbot, you know, like there's so many attack vectors, you know what I mean? And like um like you have a a feedback channel that like you know users can post feedback into and now your bot is reading it, right? And and so I think that like this the work we've put in with auto mode and you know we have a general Swiss cheese defense of like security, right? We also like yeah you know like RL against this stuff and things like that. Um I think this is really what makes cloud tag work. It like just works seamlessly with your permissions and yeah like you know we you don't want to be prompt injected in your Slack. Yeah. Yeah.

对,Dell 正在写一篇讲这个的文章。再多说一点 auto mode——这也是 Claude Tag 这么好用的原因,因为 Claude Tag 就是靠 auto mode 跑的。你可以想象,我听过很多关于 Slack 机器人'自己造还是直接买'的讨论,我的看法是:拜托,你真的不该自己造一个 AI Slack 机器人,攻击面太多了,懂我意思吧?比如你有个反馈频道,用户会往里面发反馈,而你的机器人在读这些内容,对吧?我们在 auto mode 上投入的这些工作——加上我们在安全上整体的'瑞士奶酪'式多层防御,还有针对这类问题做的 RL 训练等等——我觉得这才是 Claude Tag 真正能跑起来的原因。它跟你的权限体系无缝配合,毕竟谁也不想在自己的 Slack 里被 prompt injection,对吧。


[35:54]

Do you have any are there any more security things in the pipeline beyond that go beyond auto mode? Um I I I think I mean I I think we're very secure like like so we with claude tag you can uh peri can provision your own like sort of credentials for claude so it doesn't need to act on your uh on your behalf. You can have like claude as an identity and and that also makes it easier to audit and inspect what claude is doing.

(Simon)除了 auto mode,安全方面还有什么在管线里的东西吗?(Thariq)我觉得我们已经相当安全了。比如 Claude Tag,你可以为 Claude 单独配置一套它自己的凭证,这样它就不需要以你的身份去操作。你可以把 Claude 当成一个独立身份,这也让审计和检查 Claude 的行为变得更容易。


[36:21] Simon

Well I guess cuz claude tag is influenced by anyone who can talk to it. So it's it's got a much wider pool of people who are telling it what to do.

我想是因为 Claude Tag 会受到任何能跟它说话的人的影响——能对它下指令的人的范围要大得多。


[36:28]

That's right. Yeah. Yeah. And of course we have probes as well like with like mythos and uh sorry with fable um and that's also like a downstream effect of our safety and research work. And I think this is the moment where you sort of see AI like Anthropic being an AI safety company really paying off when you like you know we really want Claude to be able to run in an aligned way over long periods of time and like yeah automode has to be basically flawless for this to work right and it's sort of like all downstream of our like you know our being an AI safety company. We also launch trusted devices for the remote control users out here who want to be safer. Um, and for all of our remote environments, we support uh credential injection. So if you want like quad code to be able to access data dog but you don't want quad code itself to hold the data dog credential you can set up um our identity uh credential management system so that the data dog credentials are only usable by the agent but not accessible by the agent. So we insert it on the fly when the agent tries to make a data dog request. This is that that token the proxying trick, right? Where the proxy knows anytime somebody calls this an API. Dog.com address with a token, replace the token with the real thing. I love that pattern. I'm seeing that in a whole bunch of places. It feels so obviously right to me. Let's talk a little bit about the human element. And you touched on this in the keynote this morning, but um a lot of people are feeling a sense of loss now that so much of what they considered to be their role in in building software is is is being subsumed by the models. Um how do you think about that? Like um how firstly how has the past year and a half thought changed the way you think about your own craft and the value that you add?

(Thariq/Cat)没错。当然我们还有探针(probes),比如在 Fable 上——这也是我们安全研究工作的下游成果。我觉得这正是 Anthropic 作为一家 AI 安全公司真正开始兑现价值的时刻:我们非常希望 Claude 能长时间以对齐的方式运行,而 auto mode 基本上必须做到无懈可击才行——这一切都源自我们是一家 AI 安全公司。我们还为在座用远程控制的用户推出了 trusted devices,让大家更安全。另外我们所有的远程环境都支持凭证注入(credential injection):如果你想让 Claude Code 能访问 Datadog,但又不想让 Claude Code 自己持有 Datadog 的凭证,你可以配置我们的身份凭证管理系统,让凭证'可被 agent 使用、但不可被 agent 读取'——当 agent 发起 Datadog 请求时,我们在链路上实时把真凭证插进去。(Simon)这就是那个 token 代理的技巧对吧?代理层看到有人带着占位 token 调用 datadog.com 的 API,就把它替换成真正的凭证。我太喜欢这个模式了,最近在好多地方都看到它,感觉它对得不能再对了。我们聊聊人的因素吧。你今天早上的 keynote 也提到了这一点——现在很多人有一种失落感,因为他们曾经认为属于自己角色的那部分软件构建工作,正在被模型接管。你们怎么看这件事?首先,过去一年半有没有改变你对自己手艺、对自己所提供价值的看法?


[38:19] Thariq

Yeah, I I think for me and I think Kat is always such a good reminder. Ken and Boris are such good reminders of like you have to be more ambitious. They're always like, you know, you have to like like, you know, we're growing so fast, we have to be on the edge, we have to do like the best work we can. Um, I think that that's kind of like a constant reminder for me where I'm like anytime I'm like kind of like slow on something. I'm like, okay, can I do it faster? Can I be more ambitious here? Um, I think the like point on and like oftentimes the answer is claude because claude is getting better as you go. So, I'm like, "Oh, the last time I tried this, I it was with the previous model or something." Um, I think with your point on loss, I think this is real. And I I I do feel that like if you're only trying to do the same work you were doing before LLMs and now it's like a prompt, it it is like I think kind of a sad feeling. And I think the way you offset that is by being more ambitious, right? And you're like, you know, like I love like I think Jared is such a good example where he's like hand wrote all of the Zig code in his Oakland apartment in like a year barely left his house, you know, and then uh and he had so much fun doing that and now I see him like rewrite all of Bon into Rust and he's having so much fun doing that, right? and it's like so much more ambitious and that's how how like he sort of like offsets that and I think just generally being like okay like how do I do the bigger thing and and and do more and I think success is fun you know and that's how I kind of like

有的。对我来说,Cat 一直是个很好的提醒,Ken 和 Boris 也是——他们总在说:你得更有野心。我们增长这么快,必须站在最前沿,必须做出我们能做的最好的工作。这对我是一种持续的鞭策:每当我在某件事上进展慢了,我就会问自己,能不能更快?能不能在这里更有野心一点?而且很多时候答案就是 Claude,因为 Claude 一直在变强——我会想'哦,上次我试这个的时候还是上一代模型呢'。至于你说的失落感,我觉得是真实存在的。如果你只是想做和 LLM 出现之前一模一样的工作,而现在它变成了一句 prompt 的事,那确实挺让人难过的。而化解它的方式就是变得更有野心。我特别喜欢 Jared 这个例子:他在奥克兰的公寓里花了大概一年时间手写了全部 Zig 代码,几乎不出门,而且他做得特别开心;现在我看到他在把整个 Bun 用 Rust 重写,同样乐在其中——这件事的野心大得多,他就是这样化解那种失落感的。总的来说就是问自己:我怎么去做那件更大的事、做得更多?成功本身是有乐趣的,我大概就是这么想的。


[39:41] Simon

the it's changing your ambition it's changing what you do because what you did before is a lot easier in quotes but now we can take on these bigger challenges

所以是改变你的野心,改变你做的事——因为你以前做的事现在(打引号地说)'容易'了很多,于是我们就能去接更大的挑战。


[39:51] Thariq

yeah I think there's just like on average everyone has things they wish they did you know what I mean and that they were better at and I think now it's Let's let's do it, you know, like let's Yeah.

对。我觉得平均而言,每个人心里都有一些一直想做、希望自己更擅长的事,懂我意思吧?而现在就是——那就去做吧!


[40:00] Simon

And Cap, what does that look like from a sort of product management perspective?

那 Cat,从产品管理的视角看,这又是什么样的?


[40:04] Cat

I feel like the product role just changes every single month and it's very much just identifying, okay, what are um like all the PMs on our team are like this like mix of engineer, designer, PM. Um most of the engineers on our team actually used to be uh full-time engineers in the past. And so for us it really means like plugging in whenever there's any kind of gap. So if it's like okay we have this idea and we didn't inspire any engineer to go build it then like we should just build it and then put it into a notebook and inspire people to take this to production. Or if the designs look a little off, well let's let's take a page that's similar and do a first pass design and uh tag in tag in someone who's very detail oriented to like fill in the gaps. or if we notice that like hey now now our team uh and our product adoption is a bit bigger within the company and more people need to know what's coming down the pipe for quad code quad tag and co-work uh what we do then is like okay let let us like automate figuring out our whole launch calendar let's automate getting those status updates asynchronously so we're not bugging people and then let's figure out okay these are our three internal announce channels and make sure that our updates there are fully detailed and to the point. And so for us, it's very much just understanding what is the gap right now between a great idea and getting something to our customers and then how do we automate it as much as possible.

我感觉产品这个角色几乎每个月都在变,核心就是不断识别:现在的缺口在哪里。我们团队所有 PM 都是工程师、设计师、PM 的混合体,团队里大多数工程师过去也都是全职工程师出身。所以对我们来说,产品工作就是哪里有缺口就补哪里。比如我们有个想法,但没能激发任何工程师去做,那我们就自己把它做出来,放进一个 notebook 里,去激励别人把它推向生产环境。如果设计看着有点不对劲,那我们就找一个类似的页面先做一版初稿设计,再拉一个特别注重细节的同事来补齐剩下的部分。又比如我们注意到,团队和产品在公司内部的采用规模变大了,更多人需要知道 Claude Code、Claude Tag 和 Cowork 接下来有什么计划,那我们就把整个发布日历的梳理自动化,把状态更新改成异步自动收集、不去打扰别人,然后确定好三个内部公告频道,确保那里的更新详尽又切中要点。所以对我们来说,产品工作就是搞清楚:从一个好想法到把东西交付给客户之间,此刻的缺口是什么,然后尽可能把它自动化。


[41:36] Simon

It sounds to me like with product management, there's always more to do, right? It's I feel like one of the things that makes me feel good is I've never worked at a company that didn't have a backlog of a thousand things they wanted to do and didn't have the resources to take on. Um, so what's a moment when Claude has surprised you? Like when when when the model has done something that genuinely surprised you that you you didn't think it would be able to do that.

在我听来,产品管理这活儿永远有干不完的事,对吧?有件事让我挺宽慰的:我从没待过哪家公司不是积压着一千件想做却没资源做的事。那么说说,有没有哪个瞬间 Claude 真正让你吃惊?就是模型做到了某件你原本觉得它不可能做到的事。


[41:59] Thariq

Yeah, I mean I I've posted a lot about uh cloud video editing, but like like most recently I I gave a talk at the ACM Agentic conference and uh I was like, "Hey guys, do you have the edited video? I'd love to post it and share with my coms team." And they're like, "Oh, it's taking so long." And I'm like, "Okay, could you send me the raw files?" So they send me the video of me talking on stage, the video of the deck and the audio file, and they're like, "Good luck." And and so I like give this to Claude. I give it my HTML deck as well, and I'm like, "Hey, can you just like edit this together, you know?" And um what it does is like honestly incredible. Like I I'm ready to ship it. Like so it transcribes the entire video. Um it notices that sometimes the video of my deck is a little bit weird. there's like a popup of like an auto update in the middle and it's like, "Oh, I probably shouldn't use the video of your deck. Actually, what I'm going to do is I'm going to slice up and figure out which um slide you're on and instead use your h the HTML source, right?" And so, it's displaying the HTML source. Then, it's got a video of me, but you know, like I'm only taking up a small part of the stage. And so, it's cropping dynamically where I am on the stage. Um, and like I'm I'm pacing. So, it's like tracking me as I'm pacing. And I've got like this crop of me, the the deck, and then like it's transcribing what I'm saying.

有啊。我发过很多关于用 Claude 剪视频的帖子。最近的一次是我在 ACM Agentic 大会上做了个演讲,我问主办方:'剪好的视频有了吗?我想发出来分享给我们传播团队。'他们说:'哎呀,还要很久。'我说:'那能把原始素材发我吗?'于是他们把我在台上讲话的视频、拍幻灯片的视频和音频文件发给我,还说了句'祝你好运'。我把这些丢给 Claude,把我的 HTML 幻灯片也给它,说:'嘿,你能帮我把这些剪到一起吗?'它做出来的东西真的令人难以置信,我都可以直接发布了。它把整段视频转录出来;它注意到拍幻灯片的视频有时有点问题——中间弹出过一个自动更新的弹窗——于是它说:'我大概不该用你幻灯片的录像,我要做的是把视频切片、判断你当前在哪一页,然后直接用你的 HTML 源文件来展示。'所以它渲染的是 HTML 源。然后是我的视频——我只占了舞台的一小块,它就动态裁切、跟着我在舞台上的位置走。我还来回踱步,它就一路跟踪我。最后成品里有我的特写裁切、幻灯片,还有我说话内容的实时字幕。


[43:16] Simon

This was Fable, right?

这是用 Fable 做的,对吧?


[43:16] Thariq

This Fable. Yeah. Yeah. Absolutely. Yeah.

对,Fable。没错,绝对是。


[43:19] Thariq

Um and it's just like like, you know, it was a good prompt, but it was a oneshot prompt. And and then I asked it to like add some like interesting animations and graphics, and it I was just like blown away kind of, you know, and it just like does all this stuff. It does ffmpeg, it does remotion, it does.

而且,怎么说呢——prompt 确实写得不错,但那是一发入魂的 one-shot prompt。之后我又让它加一些有意思的动画和图形效果,它也做到了,我整个人都被震住了。它什么都上:ffmpeg 用上了,Remotion 也用上了。


[43:35] Simon

So I have to ask the follow-up. What can't it do? What are the things where you're still disappointed? You're waiting for Claude Fable 6 to to figure out for you.

那我必须追问一句:它还有什么做不到的?哪些地方仍然让你失望,让你在等 Claude Fable 6 来解决?


[43:44] Cat

I wanted to have better design and UX taste.

我希望它的设计和 UX 品味能更好。


[43:47] Simon

Uhhuh.

嗯哼。


[43:48] Cat

Like I feel like it's now at the point where if I give it if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But, you know, the the paddings might be off or like the the interface is it's just not delightful yet. Um, I think it kind of leans on like existing best practices for apps for how apps are designed, but I feel like for Frontier AI products, there's so many new interaction experiences that we still have to have yet to design.

现在它已经到了这样的水平:只要我写一个带详细 spec 的 prompt、说清楚我想让某个功能怎么表现,它通常都能照做。但是,padding 可能不对,或者界面就是——还称不上让人愉悦。我觉得它比较依赖现有的应用设计最佳实践,可对于前沿 AI 产品来说,还有太多全新的交互体验等着我们去设计。


[44:22] Simon

There's an Opus aesthetic. You can look at something and go, "Yeah, that was designed by Opus." It' be good if we could move beyond that.

确实存在一种'Opus 审美'——你看一眼某个东西就能说:'嗯,这是 Opus 设计的。'要是我们能超越这一点就好了。


[44:28]

Yeah. Yeah. It like I I'm very excited for future models to hopefully be like interaction design thought partners. Hm. Uh what can't it do? Um I I think I would love to see it, you know, interact more with the real world like okay like can it do this like you know can it solve science right? Like can it like orchestrate you know the experiments and there's some amount of coding that goes into that but there's also this like other taste of uh you know like the broader world that it needs. So

对,对。我特别期待未来的模型能成为交互设计上的思考伙伴。嗯,说到它还做不到什么——我很想看到它更多地和真实世界互动,比如说,它能不能做科学研究?能不能去编排、调度实验?这里面当然有一部分是写代码,但还需要另一种对更广阔世界的「品味」和理解。


[44:54] Simon

Claude Science is a new product that just came out a few days ago right?

Claude Science 是几天前刚发布的新产品,对吧?


[44:58]

Yeah but I have no context on it.

对,不过我对它没什么了解。


[44:59] Simon

I was going to ask is that part of Claude code or is that a separate separate sex section?

我正想问,那是 Claude Code 的一部分,还是一个独立的板块?


[45:03]

It's our partner team.

那是我们兄弟团队做的。


[45:04]

Gotcha. Try it out, though.

明白了。不过你们可以去试试。


[45:07] Simon

Um, so we've got I've got a couple of closing questions. Um, which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools that other companies should steal? What are the cultural hacks that people should be should be adopting from you?

好,我还有几个收尾的问题。你们觉得 Anthropic 的公司文化里,有哪些部分特别帮助 Anthropic 把这些工具用得高效,是其他公司应该「偷师」的?有什么文化上的小窍门值得大家借鉴?


[45:23] Thariq

I'll share one and then you go. Um, I'll share one for quag. So claw tag works best when you have it in a public channel and when most of your channels are public. Claw tag is able to search across all public channels to get as much context as possible to give you the highest accuracy answer. And it's only able to do this if it has access to everything.

我先说一个,然后你来。我说一个跟 Claude Tag(Slack 里的 Claude)有关的:Claude Tag 在公开频道里、而且你们大部分频道都是公开的时候效果最好。它能搜索所有公开频道,拿到尽可能多的上下文,从而给出最准确的回答——而这只有在它能访问一切的前提下才做得到。


[45:48] Cat

Yeah, I mentioned this in my keynote, but I I think I it's so important to me. I want to reemphasize like I I think the co-founders say like we don't negotiate against ourselves, you know, and I think this is really important where you're like you can imagine trade-offs in your head and talk yourself out of doing something ambitious, you know? Um or you can just try and do the ambitious thing. And I think that like we're just so often being like, okay, what if we just did it? Like what if like, you know, like is this a real trade-off or not, right? Or like and if so, like why? like where's the proof that it's a real trade-off and not just like it sounds reasonable, right? And so I think just yeah like you know make the trade-offs show themselves to you. Be as ambitious as you can.

对,我在 keynote 里提过这一点,但它对我太重要了,我想再强调一遍:我们的联合创始人常说「不要跟自己谈判」(we don't negotiate against ourselves)。这一点真的很关键——你完全可以在脑子里想象出各种取舍,然后说服自己放弃做有野心的事情;或者,你也可以直接去干那件有野心的事。我们经常就是这样:好,那我们就直接做了会怎样?这到底是不是一个真实存在的取舍?如果是,为什么?证据在哪里?还是说它只是「听起来合理」而已?所以,让取舍自己显形,尽你所能地保持野心。


[46:29] Simon

That's so because that goes against I've got 25 years of software experience that says the default answer should be no. Everything is a trade-off. Everything has a cost. And now we're having to reimagine all of those intuitions. It's it's kind of fascinating. Okay. And so final question for both of you. What is something what what's one of your favorite absurd things that you've built with Claude just because you could build it?

这太有意思了,因为它和我 25 年软件经验形成的直觉完全相反——我的直觉是默认答案应该是「不」,凡事皆有取舍,凡事皆有成本。而现在我们不得不重新审视所有这些直觉,这挺令人着迷的。好,最后一个问题问你们俩:你们用 Claude 做过的最喜欢的「纯属因为能做所以就做了」的荒诞项目是什么?


[46:54] Thariq

I can go. Well, you think um I'm working on a 2D Street Fighter fighting game uh with me as a character and like my friends as well. Um and it uses cloud code to prompt you know Gemini and honestly the sea dance model is pretty good like uh to make like video animations. Um, and uh, it works great like like it's so good at prompting. It's like, you know, it can verify like the frames to check if this was a good animation.

我先来。我在做一个 2D 的《街霸》式格斗游戏,角色是我自己,还有我的朋友们。它用 Claude Code 去给 Gemini 写 prompt——说实话 Seedance 那个模型挺不错的——来生成视频动画。效果特别好,它非常擅长写 prompt,还能逐帧验证这个动画做得好不好。


[47:21] Simon

Is this Street Fighter 2 level 2D sprites that you're generating?

你生成的是《街霸 2》那种水准的 2D 像素小人(sprite)吗?


[47:24] Thariq

Yeah, exactly. Yeah. Yeah. Yeah. Like like 2D sprites. The animation looks amazing. And it can also figure out hitboxes. It can be like, oh, you know, your fist is like here, I'll draw the JSON hitbox. Yeah. Yeah. It's like incredible. Yeah. Yeah. So, um, I don't know if I'll put this out, but it's uh

对,就是那种,2D sprite。动画看起来棒极了。它还能自己算出 hitbox(判定框),会说「哦,你的拳头在这个位置,我来把 hitbox 的 JSON 画出来」。真的不可思议。不知道我最后会不会把它发布出去,不过……


[47:39] Simon

I feel like we need a screenshot at least. This sounds I can make a screenshot happen. Yeah. Yeah.

我觉得至少得来张截图,这听起来——(Thariq:截图没问题,可以安排。)


[47:44] Cat

Mine is much more simple. Um I'm a big rock climber and a lot of my friends climb and so we have this little app that we built with quad code that where we just log all the projects that we're working on and we also uh go outdoors together a lot. So we have uh quad do all this research with workflows. Workflows is amazing. Like we brand it as a coding tool but it's amazing for doing deep research for travel. Um, I also plan our team off sites and it's good at finding venues that can fit all of us. Um, it yeah there it has a lot of uh side benefits. But anyway, um I also use workflows to just research all the climbing destinations that we might want to go to, what has direct flights from where all of us are located. It goes to mountain project and finds all the climbs that are in our grade level. It finds the Airbnb and it actually maps out like I I don't like hiking and so I care a lot about it having a very short approach. So very short walking distance from where the car parks to where the rock actually is. And so it it filters for this and so I like with existing apps I have to like manually click through mountain project but with this I just put in all of our preferences and it's just a custom app for us.

我的就简单多了。我是个重度攀岩爱好者,很多朋友也爬,所以我们用 Claude Code 搭了一个小 app,记录我们各自正在攻克的线路项目。我们也经常一起去户外,所以我让 Claude 用各种 research workflow 做调研——它真的很棒,我们把它当编程工具来宣传,但它做旅行方面的深度调研也一流。我还负责策划团队 offsite,它很擅长找能装下我们所有人的场地,附带好处一堆。总之,我也用它来调研我们可能想去的所有攀岩目的地:哪些地方从我们各自所在的城市有直飞航班;它会去 Mountain Project 上找出符合我们难度等级的所有线路;找好 Airbnb;甚至把路线都规划出来。我不喜欢徒步,所以我特别在意「接近段」要短——从停车的地方走到岩壁的距离要尽量近,它就会按这个条件筛选。用现成的 app 我得在 Mountain Project 上手动一页页点,而现在我只要把我们所有人的偏好输进去,就得到一个专属于我们的定制 app。


[48:56] Simon

So you're you're basically vibe coding Jira for mountain climbing.

所以你基本上是在 vibe coding 一个「攀岩版 Jira」。


[49:00] Simon

Exactly. That's that's pretty fantastic. Um, you know what? We have time for a couple of audience questions. Um, if you want to sh if you want to come forward and say them to me and I will repeat them so everyone can hear them. But yeah, um, tell you what, anyone who gets here first gets to ask a question. Sorry for people at the back. Hey um actually

(Cat:没错。)这真是太棒了。好,我们还有时间接几个观众提问。想提问的可以走到前面来跟我说,我会复述一遍让大家都能听到。这样吧,谁先到谁先问——后排的朋友抱歉了。


[49:21]

yeah uh my question is that um do you have any near plan to build more uh eval uh tools for us to build eval data set or anything like that and um more observability tools to monitor the performance of agents and uh workflows. We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high quality evals. And so I think the tooling is less of the constraint and more of the skill set of how do you build a great eval? And that's an area where we're excited to both invest internally and also hopefully we can share some of the best practices externally.

(观众提问)我的问题是:你们近期有没有计划做更多 eval 工具,帮我们构建 eval 数据集之类的?还有更多可观测性工具,用来监控 agent 和 workflow 的表现?(Cat 回答)我们考虑过做 eval 工具,但我认为真正的瓶颈其实在于:客户要构建真正高质量的 eval 需要很长时间。所以限制因素与其说是工具,不如说是「如何构建一个好 eval」这项技能本身。这也是我们很期待的方向——一方面在内部投入,另一方面也希望能把一些最佳实践分享给外部。


[50:04]

Hey Katar, uh my name is Sai. Uh so my question was because I'm more interested in the memory and the multiplayer. How does how is memory being designed? So two questions right uh so how is memory being designed today I I assume it's around files. So and second part of it is have you thought about thinking in an orthogonal direction where you would actually need a data store to store these memory instead of files to scale it better. So I think that's my question.

(观众提问)嗨 Cat、Thariq,我叫 Sai。我比较关心 memory 和多人协作这块,问题有两个:第一,现在 memory 是怎么设计的?我猜是基于文件的。第二,你们有没有考虑过一个正交的方向——用一个数据存储(data store)来存这些 memory,而不是用文件,以便更好地扩展?


[50:31] Thariq

Yeah right now for cloud tag the memory is channel specific. So every claude in that channel has a shared memory and then you know the instances have a session um but like the session can contribute back to main memory. We do a lot of memory research and it's you know can be kind of unintuitive like what what is the right way to do memory. Uh but yeah we're always working on this. So yeah

对,目前在 Claude Tag 里 memory 是按频道划分的:同一个频道里的所有 Claude 共享一份 memory,每个实例有自己的 session,但 session 里的内容可以回写到主 memory。我们做了大量 memory 方面的研究,这块其实挺反直觉的——到底什么才是做 memory 的正确方式。总之我们一直在持续攻关。


[50:53]

uh yeah I mean we're always running me memory experiments. I don't have you know anything to yeah like how it works right now in cloud tag is a markdown file per channel. Yeah.

对,我们一直在跑各种 memory 实验,暂时没有什么可以公布的。现在在 Claude Tag 里的实现就是每个频道一个 Markdown 文件。


[51:01]

Thank you.

谢谢。


[51:02] Simon

So I'm afraid we're out of time. Please join me in thanking Cat and Tariq and we will be around for more questions um in the hallway.

很遗憾时间到了。请和我一起感谢 Cat 和 Thariq——我们会在走廊里继续回答大家的问题。


[51:10]

Thanks guys.

谢谢大家。


[51:10]

Thank you.

谢谢。