How to Build an Agent-native Product | Mike Krieger
频道: Every
视频: https://www.youtube.com/watch?v=KRv9GpJYrUA
原文语言: en
统计: 共 38 轮 · Mike Krieger 26 · Dan Shipper 8
[0:00] Mike Krieger
The models today are good at adding features. They're not necessarily good about figuring out what to cut out of the product. You can get it to go zero not it's zero to one, but zero to end pretty quickly over the matter of hours. It's made a lot of decisions along the way and some of the sort of intuitions you build about what are the right things to to put in there. I think you build over time. I feel like that is the art and science of software design in in 2026.
今天的 model 很擅长往产品里加功能,但不太擅长判断该砍掉什么。你能让它在几个小时内就把东西从零做出来——不只是零到一,而是直接做到头。可它一路上替你做了一大堆决定,而有些直觉,比如什么东西才真正该放进去,我觉得是要靠时间慢慢积累的。这就是 2026 年软件设计里那种既是艺术、又是科学的东西。
[0:24]
[music]
[音乐]
[0:36]
The world moves fast [music] and in the age of AI, the pressure isn't just to move faster. It's to make sure that what you send actually sounds like you. From emails [music] to proposals to stakeholder updates, generic and rushed just doesn't cut it. If you've ever stared at a blank page knowing exactly what you want to say, but not how to start, Grammarly fixes that. Grammarly gives you one place to think, write, and finish your work right where you already write. Most AI tools either take over or stay out of the way. Grammarly does neither. It helps you break the blank page, adjust your tone so a message lands right [music] for the specific person reading it, and works seamlessly across more than 500,000 apps and sites that you're already using. It's loaded with agents built for [music] every step of your process and 90% of professionals say it saved them time. 93% say it helps them get more done. This is AI that works with you, not over you.
世界节奏飞快 [音乐],在 AI 时代,压力不只是要更快,而是要确保你发出去的东西真的像你自己写的。从邮件 [音乐] 到提案,再到给相关方的进度更新,套话连篇、赶工糊弄是行不通的。如果你曾对着空白页发呆,明明清楚自己想说什么,却不知道怎么开头,Grammarly 能帮你解决。Grammarly 给你一个地方思考、写作、把活儿干完,而且就在你平时写东西的地方。大多数 AI 工具要么把活儿全抢过去,要么干脆躲一边。Grammarly 两样都不是。它帮你打破空白页的僵局,调整语气,让一段话精准地打动正在读它的那个具体的人 [音乐],并且能在你已经在用的 50 多万个 app 和网站上无缝运行。它内置了为你流程里每一步打造的各种 agent [音乐],90% 的职场人说它帮自己省了时间,93% 的人说它让自己做成了更多事。这是一个与你协作、而不是凌驾于你之上的 AI。
[1:29]
[music]
[音乐]
[1:29] Dan Shipper
In a world of generic AI, don't sound like everyone else. With Grammarly, you never will. Download Grammarly for free at grammarly.com. That's grammarly.com. Mike, welcome to the show. Great to be here. Thanks for having me on. Great to have you. Um I I'm super excited. For people who don't know, you are the co-founder of Instagram and now you are at Anthropic and Anthropic Labs. Um I'm you know, I've admired your work from afar both at Anthropic and at Instagram for a really long time and you're obviously at the forefront of building products in AI. So, thank you for coming on. Absolutely. Where should we start? Like what what we were talking about just now in the in the pre-production is Um what has gotten easier and what has gotten harder or stayed maybe stayed the same in product building as we've as as the underlying substrate or the process by which we build products has changed completely. Um so, like tell me about your experience now versus, you know, earlier in Anthropic versus Instagram and and how you think things are changing. Yeah, I was doing the thought exercise a couple weeks ago of, you know, we know in the Instagram story we had another product called Bourbon. We worked on that for almost a year. Uh it wasn't working. We pivoted. We basically spent three months building what became Instagram, launched it, and then scaled it. And I was asking the question like what is now trivial um and what was actually inherent in that building process that doesn't get easier, right? And that year we probably could have hit some of the dead ends we had eventually hit sooner, but there was value in getting there too, right? Like we over complicated the product so that we then had to simplify it. I find even the models today are good at adding
在一个满是套话 AI 的世界里,别和所有人听起来一个样。有了 Grammarly,你绝不会。去 grammarly.com 免费下载 Grammarly,就是 grammarly.com。Mike,欢迎来到节目。很高兴来这儿,谢谢你请我。很高兴你能来。我特别激动。给不认识你的人介绍一下,你是 Instagram 的联合创始人,现在在 Anthropic 和 Anthropic Labs。我一直在远远地仰慕你的工作,无论是在 Anthropic 还是在 Instagram,仰慕了很久,而你显然站在用 AI 做产品的最前沿。所以谢谢你来。当然乐意。我们从哪儿聊起?刚才在开场前我们聊到的是——随着我们做产品所依赖的底层基底、或者说整个流程彻底变了,产品打造这件事里,哪些变简单了,哪些变难了,又有哪些可能没怎么变。给我讲讲你现在的体验,对比你早期在 Anthropic、再对比 Instagram 时期,以及你觉得事情正在怎么变。是这样,几周前我做了个思想实验。Instagram 的故事大家都知道,我们之前还有个产品叫 Bourbon,做了差不多一年,没做起来,我们就转型了。我们基本上花了三个月做出后来的 Instagram,上线、然后做大。我当时在想:现在哪些事情已经变得不值一提了,而当年那个打造过程里,又有哪些东西是本质上就不会变简单的?那一年里,我们其实本可以更早撞上后来终究会撞上的那些死胡同,但走到那一步本身也是有价值的,对吧?比如我们把产品搞得过于复杂,所以后来才不得不去给它做减法。我发现连今天的 model 都很擅长加功能——
[3:05] Mike Krieger
features. They're not necessarily good about figuring out what to cut out of the of the product and that took a lot of just uh sort of, you know, hitting actual actual real world usage. Um and there was something about the process of incrementally adding things right now. I mean, today especially some of the stuff we're building in labs like you can get it to go zero not it's zero to one, but zero to end pretty quickly over the matter of hours, but it's made a lot of decisions along the way and yeah, you can ask it to follow up with you and and then do input, but some of the sort of intuitions you build about what are the right things to to put in there. I think you build over time. And so, I I've been reflecting like there haven't been a lot of breakout consumer products even in the age of accelerated AI building and I think part of it is because it just still takes time to sort of hone your view about what sort of intervention you want to make on the world and then build from there. Now, the actual building part once you know what to build is of course so much easier. I I had Claude basically rebuild Bourbon. It took about two hours. It was feature complete. It added filters which Bourbon didn't have. We added those for Instagram, but I think it knew, you know, uh it knew what the eventual future of the product so it decided to build that in. Um so, I think that that part feels feels really different. But I think there's also, you know, uh I remember there was a week where Kevin went off uh and built all the filters for Instagram V1. I went off and built like sort of the rest of the app. And, you know, sitting there I was I would stay up till 4:00 a.m. and then sleep till noon. That's like my natural day-night cycle. And like in that
——但它们不太擅长判断该从产品里砍掉什么,而这一点得靠大量真实世界的使用去磨出来。还有一点是关于一点点往里加东西的那个过程。我是说,今天尤其是我们在 labs 里做的一些东西,你能让它从零做出来——不只是零到一,而是几个小时内就直接做到头——但它一路上做了一大堆决定。是的,你可以让它回头再跟你确认、再让你输入意见,但有些直觉,比如什么东西才真正该放进去,我觉得是你要靠时间积累出来的。所以我一直在反思:即便在 AI 加速开发的时代,也没冒出多少现象级的消费级产品,我觉得部分原因就是,你仍然得花时间去打磨自己对世界的看法——你到底想对这个世界做出什么样的干预——然后再从那里开始建。当然,一旦你知道要建什么,真正动手建的那部分确实轻松太多了。我让 Claude 基本上把 Bourbon 重做了一遍,大概花了两个小时,功能就齐全了。它还加了滤镜,Bourbon 当年是没有的,那是我们后来给 Instagram 加的,但我觉得它知道这个产品最终会演变成什么样,所以就主动把那个做进去了。所以我觉得那部分感觉真的很不一样。不过我也记得,当年有那么一周,Kevin 跑去把 Instagram V1 的所有滤镜都做了,我跑去把 app 的其余部分都做了。当时我会熬到凌晨四点,然后睡到中午——这就是我天然的作息周期。而在那个——
[4:26] Mike Krieger
process you're making so many decisions. Like how should location work? How do And, you know, it's we got to find a way of accelerating building while still sort of helping people build intuition of those decisions along the way cuz otherwise I think you either get just get very generic products that are unlikely to break out or ones that just don't reflect some deeper intuition that you come to about your space or your product. This is great. I love this. Um it's making me think of two things. One is um I have this like little thing in my head that if you grow a tree without it like with it being [clears throat] indoors without it being exposed to wind, it doesn't get as strong as cuz as it's growing it needs all these forces pushing it like back and forth in order to like make a a real tree. And so, if you if you have it indoors without wind, it you're going to grow a tree, but it like leans and it gets all and it's not as strong and it's not it's not the same thing. And I think there's something that you're saying here where because we've accelerated the pace of development so drastically, um what what would normally be this sort of incremental thing where you're you're doing things one at a time and then you're exposing it to users, you can actually kind of grow an entire tree indoors and then you have this like whole thing that you're just like it doesn't have the same um level of intuition and exposure to experience at each step that that creates a a great product. Is that is that is that
——过程里,你要做出特别多的决定,比如定位功能该怎么运作之类的。所以我们得想办法,在加速开发的同时,还能帮人在过程中建立起对这些决定的直觉,否则我觉得结果要么是做出特别套路、不太可能爆的产品,要么是做出那种没体现出你对自己领域或产品更深层直觉的东西。
[5:51] Dan Shipper
I love that. I I I um I love that metaphor too. We, you know, when we were starting Instagram we had this we were very into like Eric Ries and Lean Startup and that whole like YAGNI like you ain't going to need it principle. And um I have found and actually even one of the things I was working on in labs recently, we way over built for V1 before we even got to early access because you can. You're like, oh, well, we have this option. Why not add this one as well? That's like that's a PR of work. And if you get a really good flow and Claude code, you know, you're firing things off. You're going to lunch. You're coming back. The thing is done. You're like, great, we added it. And the thing we realized was we'd created this sort of matrix of functionality that was actually quite hard to test and and keep up with right before launch or even to explain to people. Like they're arriving. The metaphor I somebody else gave me which I really like is the difference between sort of getting episode by episode, getting to characters in a TV show versus imagine like you're thrown into the final episode. You're like, wait, what are all these things and who are all these people? And like I already, you know, I'm expected to have all of this context. I think there's the same kind of feeling around like developing something over time. But I the tree metaphor I think sticks too as well. And so, like showing somebody the fully formed tree is also kind of a lot all at once. I I think there's there's there's definitely something there in how do you build product these days and still keep it simple. And not because just because you can doesn't necessarily mean that it should be in at least the first version. I'm having the same problem because you know, I was I was literally up until
这太棒了,我太喜欢这个说法了。它让我想到两件事。第一件,我脑子里有个小比方:如果你养一棵树,把它养在室内、不让它经风,它长出来就不如经过风吹的那么结实,因为它在生长时需要这些力来回推它,才能长成一棵真正的树。所以如果你把它放在没有风的室内,你确实会长出一棵树,但它会歪,会各种毛病,不够结实,跟外面的树不是一回事。我觉得你刚才说的就有这层意思:因为我们把开发节奏加速得太猛了,本来那种循序渐进的过程——你一次只做一件事,然后拿给用户去试——你现在其实可以在室内就把一整棵树养出来,然后你手上一下就有了这么个完整的东西,但它没有那种在每一步都积累起来的直觉、没有经受过经验的检验,而正是这些才造就一个伟大的产品。是不是这个意思?
[7:14] Mike Krieger
4:00 a.m. debugging and and fixing this app that I made like on the side at Every called Proof which is a agent-native collaborative markdown editor. So, you can like share share really quick plan docs and stuff with your team or with other agents. And you have little presents and it's really fun. And I is this is like my second or third iteration of the full product end-to-end which is really interesting you can do now. But the first couple iterations, I just found myself because vibe coding is so fun and so addictive, I just found myself being like, yeah, like I'll do this and I'll do this. And like and it just created this monstrosity that wasn't that good to wasn't that good to use. And I got really inspired by we have another product called Monologue which I'm not sure if you've run into or not. But I got really inspired by Monologue which is a really simple speech-to-text app run by um JM Naveen who he's just so focused on making one simple thing work so well. And I saw how well that works in this age of just like anyone can make a product is like something that's super polished and just super good at what it does. And so, I just basically threw out the product and started over with this very simple like it's just a shareable markdown uh link. And that then just like started growing virally inside of every like everyone started using it all the time. And then now I we launched it and it just blew up. And so, I spent all last night like not sleeping trying to fix it. And being like, I'm too old for this [ __ ] I can't I can't be doing this anymore uh cuz it just reminded me of like being in my 20s or like being in college and like hacking on [clears throat] stuff and whatever which is fun, but also exhausting. Um and so, yeah, I've I've found that I've had to
我太喜欢了,这个比喻我也特别喜欢。我们当年创办 Instagram 的时候,特别推崇 Eric Ries 和《精益创业》,还有 YAGNI 那一整套——就是 "You ain't gonna need it"(你根本用不上)的原则。我发现——其实我最近在 labs 做的一个东西就是这样——我们在还没进早期内测之前,就为 V1 大大地超额建设了,因为你真的做得到。你会想:哦,反正我们有这个选项,干嘛不顺手把这个也加上?那不过就是一个 PR 的工作量。如果你跟 Claude code 配合得特别顺,你噼里啪啦把任务都派出去,去吃个午饭,回来一看,东西做完了,你会想:太好了,加上了。然后我们意识到,自己造出了一个功能的矩阵,结果在临上线前特别难测试、难维护,甚至难跟别人解释清楚。用户一进来就懵了。有人给我打过一个我很喜欢的比方:这就像一部电视剧,你本该一集一集地认识那些角色,但现在你被直接扔进了大结局那一集,你会想:等等,这都是些什么、这些人都是谁?而我却被默认已经掌握了所有这些背景信息。我觉得随时间慢慢开发一个东西也是同样的感觉。不过那个树的比喻我觉得也一样成立。所以把一棵已经完全长好的树一下子摆给别人看,本身也是信息量太大了。我觉得这里头确实有点东西,关于这年头到底该怎么做产品、又怎么把它保持简单。不是说你能做,就一定意味着至少在第一版里就该把它放进去。
[8:55] Dan Shipper
really modify my psychology because so much is possible. How are you dealing with that? Yeah, and just as a brief aside on that, I mean, with Bourbon, our biggest mistake was adding functionality over time rather than deleting it, right? And because oh, you know, eight features doesn't make for good product. Maybe the ninth one will. Instead it just made for, you know, something that felt really complicated. I mean, I think a couple of things are also like part of how we're dealing with it is actually being more willing to do rewrites. Um you know, like classic, you know, Fred Brooks' Mythical Man-Month. Like you you shouldn't rewrite software because all the things that were imbued in V1, you're going to mess up. And it also leads Yeah. Yeah. Exactly. And the whole second system syndrome. And there is still a lot of truth to that, but one, you know, the models can help you sort of diff and basically see did you miss anything that was in that first one. But second, it's just it's no longer you're not like talking about a year-long rewrite that might have killed a company like, you know, famous like Netscape like where these are like days probably especially off of a given source. So, we've actually had several initiatives like usually pre-launch, rarely post-launch, but at least pre-launch like have built the the thing, realized we've overcomplicated or made some kind of core assumption, Um and then like tore it down, done a V2, and then and then iterate on it from there. So, it doesn't surprise me that that's become sort of part of what you've had to do as well, but it doesn't feel as painful. You're not like, "Oh, like my year of building this thing." It's like, "Oh, that was last week, and then I got to do it this week, and I get
我也遇到了一模一样的问题,因为我昨晚真的熬到——
[10:13] Dan Shipper
to cut out a lot of a lot of what was there as well." Um I think functionality-wise and how we're dealing with it from a product development standpoint, um I think we are learning to launch earlier, um and it's definitely a balance around, you know, we've grown, we have like a strong enterprise footprint, people have expectations about like what the initial version is, um but not assuming that we're going to know what every connector, everything that we need to add to the product is but ahead of launch cuz people still will absolutely surprise us, right? We're We have a strong contingent in uh contingent of we call them ant footers cuz we're ants at Anthropic, um but that only that only gets you so far before you need that that real-world contact. Like take Cowerk for example, we'd been noodling on a product of that shape for a long time. Um and then uh once we decided like, "No, let's get this out. Let's actually, you know, build the build the V1 that we think solves the problem in the most minimal way possible and get that out in 10 days." Was really a good push around, yes, there are a hundred things that V1 should or could have had, um but it didn't. And at the same time it was it was useful enough to prove something out there, and I'm not sure developing it for another 2 months adding, you know, 50 features would have been more useful. In fact, we probably would have been building in a the indoor tree would have been getting built and then the second it hit real-world it's like, "Actually, nobody wants to do that. They want to do this this other piece." So, I think that piece to that again, there's like the intuitions of the original lean startup ideas are still here. It's just they manifest at different time scale and in a different way.
——凌晨四点,在调试、修一个我在 Every 顺手做的 app,叫 Proof,它是一个 agent-native 的协作式 markdown 编辑器。你可以很快地把计划文档之类的东西分享给你的团队或者别的 agent,上面还有小的在线状态显示,特别好玩。这已经是我把整个产品端到端做出来的第二、三个版本了,现在能这么干真的很有意思。但头几个版本,我发现自己——因为 vibe coding 实在太好玩、太上瘾了——总是忍不住想:行,我加这个、我再加那个。结果就造出了一个四不像,用起来其实没那么好。后来我受了很大启发,我们还有个产品叫 Monologue,不知道你有没有遇到过,是个特别简单的语音转文字 app,由 JM Naveen 负责,他就是极度专注于把一件简单的事做到极好。我看到在这个人人都能做产品的时代,那种打磨得特别精致、把自己那件事做到极致好用的东西有多管用。于是我基本上把那个产品整个扔了,从头来过,做了这个特别简单的东西——就是一个可分享的 markdown 链接。然后它就在 Every 内部病毒式地传开了,每个人都开始一直用它,接着我们正式上线,结果一下就爆了。所以我昨晚整夜没睡都在修它,心里想:我这把年纪不该再折腾这种破事了,真的受不了了,因为这让我想起二十几岁、或者上大学时熬夜瞎搞东西的那种感觉,好玩是好玩,但也是真累。所以是的,我发现我不得不——
[11:38] Mike Krieger
I'm really curious to hear how you think about product design and how products should work because the I I've been Anyone that ever will tell you the the the the the phrase that I use the most or the word I use the most about the software we build is it has to be agent native. Um so, agents have to be able to like use it as anything that an agent a user can do in the in the app, the the agent can do. There's a couple other like little principles of being agent native, but I basically stole that from you guys. Like I think that Claude Code is the canonical thing that taught me about how that kind of product can work so well where it's like it's an agent, it can do anything on your computer that you can do, um and it's customizable and flexible and extensible, so it's easy to start, but it can do all sorts of unexpected things that um the designers didn't really like think about beforehand. And I think that that's such a good model for AI product development in AI, and I'm kind of curious like this is just sort of what I've cribbed from watching what you guys do and then like kind of put my own spin on but how do how do you think about it and how do you how do you talk about making products like that? Yeah, there's so much in here, and I love the agent native write-up you all did. It's like to me the canonical exploration of this. So, thanks for like putting that ideas out in a in a really in a really clear way. So, I think a few threads to pull on this. One is a conversation I had with somebody recently where they said, you know, like you all they're a non-technical person. They're like, "You all are talking about like agents and all this stuff." Like they're just like, "Actually, computers just work now. I always wanted computers to work and they
——好好调整一下自己的心态,因为现在能做的事太多了。你是怎么应对这种状况的?是这样,先简短插一句:当年做 Bourbon,我们最大的错误就是随着时间不断往里加功能,而不是删功能。因为我们会想:哦,八个功能撑不起一个好产品,也许第九个就行了。结果它只是变成了一个让人感觉特别复杂的东西。我觉得还有几点,比如我们应对的方式之一,其实是更愿意去做重写。你知道那个经典的 Fred Brooks《人月神话》——它说你不该重写软件,因为 V1 里融进去的那些东西,你重写时都会搞砸,而且还会导致——是的,没错,就是那个 "第二系统综合症"。这话到现在仍有很大道理,但第一,model 能帮你做 diff,基本上看出你有没有漏掉第一版里的东西。第二,现在这已经不再是动辄一年的重写、那种可能把公司搞垮的重写了——比如著名的 Netscape——现在大概就是几天的事,尤其是在已有一份源码的基础上。所以我们其实搞过好几次这样的行动,通常是在上线前,上线后很少,但至少在上线前:把东西先做出来,发现自己搞得过于复杂了、或者做了某个核心假设是错的,然后把它推倒,做个 V2,再从那里往下迭代。所以你也变成不得不这么干,我一点都不意外,只不过它不再那么痛了。你不会再想:天哪,我这一整年都搭进这东西了。而是:哦,那是上周的事,这周我又得重做一遍,而且我还能——
[13:13] Mike Krieger
didn't work and now they work." And it's just a funny thing where if you knew the incantations to properly get on the command line and brew install the thing that like nobody is going to do that, but now Claude can do it for you and therefore like the computer now feels like a a tool that is alongside you, and I think that's that that core insight is it's more than even just adding power and functionality to new software. It's also just unlocking the functionality that always should have been there or available and just felt like extremely hard for people. So, that's like maybe thought number one. Thought two is actually comparing our products that do this well and versus not. I think Claude Code does it well. I think Claude AI still needs to evolve a lot. So, um as an example, I was watching somebody use Claude and they were in a project and they had built I think an artifact or or a new document. They said, "Great, can you add this to my project knowledge?" And Claude's like, "Yeah, let me tell you the steps to go add it to my project knowledge." Like, "No, that should just be a thing that it can do really natively." And so, I think even in that you see a product that was a 2024 product that has been iterated on and evolved a lot, but still I don't think has been baked in from the very beginning the idea that every single one of its primitives it should have knowledge about and the ability to modify. And I think that's essential in products these days. And I think Claude Code is the 2025 vintage of that, and I think there's even further aspects of it when you see what some of the the harnesses that folks are experimenting with where they can actually sort of modify the harness itself. That starts getting to the next maybe level of that
——顺手砍掉之前一大堆东西。我觉得在功能层面、以及我们从产品开发角度怎么应对这件事上,我们正在学着更早地上线。这里头肯定有个平衡:我们公司已经做大了,有相当强的企业客户基础,大家对初版该是什么样子是有预期的;但我们也不再假设自己在上线之前就能想清楚每一个 connector、产品需要加的每一样东西,因为用户绝对还是会给我们惊喜的,对吧?我们内部有一支很强的力量,我们管他们叫 "ant footers"(蚂蚁内测员),因为我们在 Anthropic 自称是蚂蚁,但这只能帮你走到一定程度,再往后你就需要真实世界的接触了。拿 Cowork 举例,这种形态的产品我们琢磨了很久。后来我们一拍板:不,把它放出去,咱们就把那个我们认为能以最精简方式解决问题的 V1 真正做出来,10 天内放出去。这是一次很好的推动——是的,V1 本该或本可以有上百样东西,但它没有。可与此同时,它已经够好用,足以在外面验证出点什么了,我也不确定再多开发两个月、加上 50 个功能会不会更有用。事实上,我们多半又会在搭那棵室内的树,结果它一接触真实世界就发现:其实根本没人想要那个,大家想要的是另外那块。所以那部分我觉得,精益创业那些最初的直觉到今天依然在,只是它们以不同的时间尺度、用不同的方式显现出来了。
[14:35] Mike Krieger
where, you know, it's probably esoteric for most people, but even unlocking that functionality means that you don't have to um sit there and be like, "Oh, I I wish it did this a little bit differently, you know, I wish Gmail worked in a slightly different way." Instead of just asking it to. And and I think that that feels like the big next step, but um even within like Claude Code just teaching Claude Code about Claude Code was a a really valuable experience. I was like this definitely relates. This is now getting very circular meta, but bear with me. Uh I loved your write-up on agent native. I was like, "I want this as a skill." So, whenever I'm prototyping something, it thinks in an agent native way. So, I had it packaged it up as a skill, and that whole process was, you know, "Hey, Claude and Claude Code, I, you know, can you create a skill for this?" It's like, "Sure, I'm looking up my skill skill. I'm going to create a skill about it. I'm going to install it." I'm like, "Great, is that available now or do I need to reload?" It said, "All right, I think you need to restart it. Let me check. Yep, you do. All right, let's go." And everything was it has knowledge about itself, and that unlocks so much uh capability in there as well, which maybe is like the last thread to pull on. I think all of these could be hour-long conversations, which is I think and one of the things that we're really thinking about in labs is how do you imbue the software that Claude builds to be more Claude aware and even just Claude agent native sort of building aware so that it even thinks to build in that way to start with cuz it still won't partially because uh decades of software is not that, right? So, how do you get new software to have that principle baked in?
这功能对大多数人来说可能有点小众,但哪怕只是把它解锁出来,就意味着你不用干坐在那儿想:「唉,我真希望它的行为能稍微改一下,我真希望 Gmail 能换一种方式工作。」而是直接让它去改就行了。我觉得这是很大的下一步。不过哪怕是在 Claude Code 里面,光是教 Claude Code 了解它自己,就已经是一段非常有价值的经历。我当时就想,这肯定有关联。现在这话题越来越绕、越来越元(meta)了,但你听我说完。我特别喜欢你那篇关于 agent-native(智能体原生)的文章,我当时就想:我要把这个做成一个 skill。这样我每次做原型的时候,它都会用 agent-native 的方式来思考。于是我就把它打包成了一个 skill,整个过程就是:「嘿,Claude、Claude Code,你能帮我做一个这方面的 skill 吗?」它说:「当然,我正在查我那个『做 skill 的 skill』,我要做一个关于这个的 skill,然后把它装上。」我说:「太好了,现在能用了吗,还是我得重新加载一下?」它说:「好,我觉得你得重启一下。我查一下。对,是要重启。好,走起。」整个过程里,它对自己是有认知的,这又在里面解锁了非常多的能力,这可能就是最后一根可以拉一拉的线头了。我觉得这些话题每一个都能聊上一个小时。我们在 Labs 真正在琢磨的事情之一,就是怎么让 Claude 构建出来的软件更「懂 Claude」,甚至天生就带着 agent-native 的构建意识,让它从一开始就会想着用那种方式去构建。因为它现在还做不到,部分原因是几十年积累下来的软件根本不是这样的,对吧?所以问题是:怎么让新软件从根上就把这个原则烙进去?
[15:58] Mike Krieger
That's the thing I was about to ask you about. Like so, A, I'm super I I'm super honored that you're you that you read the write-up and you're using you made a skill for it. That's that's amazing. And B, like yeah, you're pointing to a real problem that I have found is um I think actually Claude models are the best for this. Like a Codex model generally is not as good. Um uh at building an agent native because they're uh models in general unless you push them, they think like traditional engineers. And that's a whole different set of you know, you want to have guardrails and tests. You want to make sure that there's like one path the user can go down versus for creating this extensible thing that's super flexible. Um So, yeah, how how are you how do you how do you architect your product to teach the models and the harness to teach teach the models to think and work in this way? Yeah, I think there's two parts to it. One is the more sort of mundane part, and the second one I think is the one that's more sort of interesting in developing. The first one is like even just having good patterns and paradigms available to the model while it builds has been really valuable. Finding the right balance of templatized to skillified, right? And like what that what that right balance is. But having, you know, one of the things that we like have now is a skill about the Claude API, which sounds super obvious, but even just having that is really valuable because you would sometimes find, you know, we'd launch a new model. It wasn't in the the model's sort of innate knowledge, and then you'd get into these really funny arguments. Like, "No, I know you made a typo. It's it's Sonnet 45." You're like, "No, I I know it's Sonnet 45." And you're like, "No, no, no." So, um
这正是我刚才想问你的事。首先,我特别荣幸你居然读了那篇文章,还专门为它做了个 skill,太棒了。其次,你点到了一个我确实遇到过的真问题:我其实觉得 Claude 模型在这件事上是最强的,像 Codex 那类模型一般就没这么好。模型总体上来说,你要是不推它一把,它就会像传统工程师那样思考。而那是一整套完全不同的思路——你会想要有护栏、有测试,想确保用户只有一条路可走;而要做出这种可扩展、超级灵活的东西,要的是另一回事。所以,那你是怎么——你是怎么去设计你的产品,来教模型、教 harness、教模型用这种方式去思考和工作的?我觉得这里面有两部分。一部分比较平淡无奇,另一部分我觉得更有意思、也更值得展开。第一部分是:光是在模型构建的时候给它准备好优秀的范式和模式,就已经非常有价值了。关键是找到「模板化」和「skill 化」之间的平衡点,对吧?以及那个平衡点到底在哪。我们现在有的一样东西,就是一个关于 Claude API 的 skill,这听起来超级理所当然,但光是有这个就特别值钱。因为你有时候会发现,我们一发布新模型,它就不在模型自身的固有知识里,然后你就会跟它陷入那种特别好笑的争论:「不,我知道你打错字了,那叫 Sonnet 45。」你说:「不,我知道就是 Sonnet 45。」它还说:「不不不。」所以——
[17:32] Mike Krieger
like like having that capability, having like good templatized examples of that in skills I think helps. But then the second part is what's also interesting is that class of software is just a different type of test. Like it's much harder to sort of write an end-to-end functional test around an agent native product because part of it is that unpredictability. And so, another idea we've been kicking around a lot in labs is like how do you increase like the sort of fidelity of the verification? The other day I had a agent native iOS app that I was working on, and I was I was having Claude interact with it, and it Claude was ended up having a conversation with itself in like a chat feature in the iOS. It was very funny watching Claude talk to Claude cuz it's like somebody's pretending to like be what humans are. And this particular one was a prototype I was doing about like a sort of like work journal reflections, and the Claude was like, "Yeah, my boss is really rough on me. Like I had a hard day." And then the Claude's like, "Oh, I'm so sorry to hear that." And they're just going back and forth. But you wouldn't have written a unit test for this, and you know, maybe it would have come up with some other emergent idea as well. So, yeah, I think you just have to go much more towards, um you know, setting up harnesses that are actually exercising as much of that agent native capability as possible because you don't exactly know what things are going to do. And things are going to end up in a weird place where Claude's going to try to do something that you wouldn't even think it was going to um do, and it might put your app in a in a new state. So, maybe it's circling all the way back to still like what's hard. It's like um having the underlying architecture to
所以说,有这种能力、在 skill 里有这类好的模板化示例,我觉得是有帮助的。但第二部分同样有意思的地方在于:这一类软件就是一种完全不同类型的测试。你很难围绕一个 agent-native 产品去写端到端的功能测试,因为它本身就带着那种不可预测性。所以我们在 Labs 一直反复琢磨的另一个想法是:怎么提高验证的「保真度」?前两天我在弄一个 agent-native 的 iOS app,我让 Claude 去跟它交互,结果 Claude 在那个 iOS 的聊天功能里跟自己聊了起来。看 Claude 跟 Claude 对话特别好笑,因为它像是在扮演人类的样子。我做的这个原型大概是个「工作日志、反思」类的东西,于是一个 Claude 说:「唉,我老板对我特别凶,我今天过得很难。」另一个 Claude 就说:「哦,听到这个我真替你难过。」它俩就这么你来我往地聊。但你是不会为这种情况写单元测试的,而且说不定它还会冒出别的什么涌现出来的点子。所以我觉得,你得更多地往那个方向走——去搭建一些 harness,尽可能多地去「演练」那些 agent-native 的能力,因为你并不确切知道它会做出什么事。事情会落到一些奇怪的境地,Claude 会去尝试做一些你压根想不到它会做的事,它可能会把你的 app 带进一个全新的状态。所以这也许又绕回到了「什么是难的」这个问题上:就是要有底层架构来——
[18:55] Mike Krieger
still be robust to that is really important, right? It's like it's agent native, but it's also able to flex in a way that you might not have anticipated, but you've got the right primitives, right? I feel like that is the art and science of software design in in 2026. That's really interesting. I I totally agree with you. Yeah, you wanted to have a playground within a safe a safe environment. That's the only way you can have a playground is if it's safe around the edges, but I think initially we we made the playground like way too small and constrained. And now the models have changed, and so we can open it up a lot, but we still haven't figured out exactly like at least I I have not figured out exactly what the lines are. Um Yeah, I think that there's there's there's so much here. Like one thing that this is making me think of is that I have this idea in the back of my head, and I'm I'm wondering if you're if you if you have a a way to put this that that is more succinct. It's like the unit of value in in products right now is um it's it's like proof of work or proof of use where when someone on the team submits a PR to me, I want to see not necessarily that all the tests passed cuz I just assume that it did, but like send me a loom of you using it or your agent using it so I can tell is this good or not, you know? Yes. Yeah, how are you how are you thinking about that? Yeah, I think that there's probably like three layers to that. There's like the first one was like Claude prove to me that you've exercised this in some way, you know? I've started doing that in all my promises. I end you know, when it's working on a feature I'm like and by the end, you know, before you PR like prove to yourself and then to me that it works as intended. Like find the right way of
——来在面对这种情况时依然保持稳健,这一点非常重要,对吧?它是 agent-native 的,但同时又能以你可能没预料到的方式灵活伸缩,因为你已经有了对的基本组件(primitives),对吧?我觉得这就是 2026 年软件设计的艺术与科学。这真的很有意思。我完全同意你的看法。对,你想要在一个安全的环境里有一个「游乐场」。能有游乐场的唯一前提,就是它的边缘是安全的。但我觉得我们一开始把这个游乐场做得太小、太受限了。现在模型变了,所以我们可以把它放开很多,但我们还是没完全搞清楚——至少我还没搞清楚——那些边界线到底在哪。对,我觉得这里面有太多东西了。它让我想到一件事:我脑子后面一直有个想法,我想问问你有没有一种更简洁的说法。就是——现在产品里的「价值单元」其实是一种「工作量证明」或者说「使用证明」:当团队里有人给我提一个 PR,我想看到的不一定是所有测试都通过——因为那个我默认它过了——而是给我发一段你在用它、或者你的 agent 在用它的录屏(loom),这样我才能判断这东西到底好不好,你懂吧?是的。对,你是怎么想这件事的?我觉得这大概有三层。第一层就是:Claude,你得向我证明你以某种方式实际演练过这个东西,对吧?我现在在所有的 prompt 里都开始这么干了——当它在做一个功能的时候,我会说:到最后、在你提 PR 之前,先向你自己、再向我证明它确实按预期工作。也就是说,找到合适的方式去——
[20:30] Mike Krieger
doing it. Which actually ends up you have to change your own sort of way you build and scaffold around saying what is the right way to get Claude able to at least test this change, you know, succinctly rather than what it likes to do is like I read the code it looks good. I'm like you wrote the code. I don't trust you. So you know, you got to really test this thing. Um and then the second one is that what you described is like, you know, everything having some, you know, sort of proof around like did is it is it working as intended and as you intended to because Claude is going to make or any of these models is going to make a lot of decisions um for you and sometimes you don't you know, I'll have engineers on the team uh put up a PR and I'm like oh, why did you choose to do this versus that? And many times the answer is they didn't choose. It was just the choice the model made and maybe it was a reasonable choice. It was probably a reasonable-ish choice, but it wasn't like the optimal choice as it fit into the paradigm. I feel like that is the uh it's it's not just proof of work, but it's like proof of thoughtfulness. Like did you think this through? Um and uh I was talking to an engineer yesterday and they're they're they're was like oh, I was really I knew you were going to ask me a lot of questions about this. So I was reviewing what Claude had done so that I wouldn't be like uh I'm not sure. You know, that's I don't I don't push on that for most PRs, but when there was one that's like oh, I'm refactoring this system and there's going to be these new primitives like great. Let's make sure those are good and that you've thought through how they interrelate cuz uh it's very easy to end up otherwise with sort of this tower of assumptions that you're not
——去做这件事。而这其实最后会逼着你改变自己构建和搭脚手架的方式,去想:什么才是让 Claude 至少能简洁地测试这个改动的正确方法?而不是任由它去做它爱做的那套——「我读了代码,看起来没问题。」我说:代码就是你写的,我才不信你呢。所以你必须真的去把这东西测一遍。然后第二层,就是你刚才描述的那个:让每件事都带上某种「证明」——它是不是在按预期、按你的意图工作。因为 Claude,或者说任意这些模型,会替你做出大量决策,而有时候你并不——我会碰到团队里的工程师提一个 PR,我问:诶,你为什么选了这种做法而不是那种?很多时候答案是:他们没选,那就是模型做的选择。也许那是个合理的选择,大概率是个「还算合理」的选择,但它并不是放进整个范式里最优的那个选择。我觉得这就是——它不只是「工作量证明」,更是「用心程度的证明」:你到底有没有把这事想透?昨天我跟一个工程师聊,他当时是这么说的:哦,我知道你会就这个问一堆问题,所以我提前把 Claude 做过的东西都过了一遍,免得到时候我支支吾吾「我也不太确定」。对大多数 PR,我其实不会这么逼问;但碰到那种「我在重构这个系统、会引入这些新的 primitives」的 PR,那就太好了——咱们一定得确保这些东西是好的,确保你想清楚了它们之间是怎么相互关联的。因为不这样的话,很容易最后就堆出一座你自己都没完全意识到的「假设之塔」——
[21:52] Dan Shipper
fully aware of. I had literally the same experience today because um I I made proof. Totally vibe coded and it's growing really fast right now, but it's going down a lot. And so I've been spending the last 12 hours like trying to fix it and so we have a little SWAT team internally at Every that like signed up to help me fix it. And so I had to like onboard them. And I was like [ __ ] How do I explain how this code base works? And so I had to like go back and forth with the model a bunch to be like okay, help me to like define these terms. Help me to like figure out how to how I can explain this so I don't look like a total idiot because like yeah, there's I understand some of it, but not all of it. Definitely not enough to like the way that I would used to have to to know to know. Um and it's a whole different thing to be like do I need to know that anymore? Is it like where's the line now? It's hard hard to tell. Which maybe get to something else and I haven't tried articulate this so bear with me as I like, you know, kind of get there which is there's products that you use that feel robust underneath and those ones that you use that you're like it feels like it's one wrong command or click away from the whole thing either like freezing or being slow. For us at Instagram like we had um uh Instagram view direct messaging view one and that like who knows? You send a message it might or may not arrive to the other person. Like we'd like wrote I wrote our own like bespoke real-time system. It was like uh you know, fell over a bunch of times and you would not trust that to send a message that you really needed somebody else to see. It was just a you know, more of a social thing. And when we built V2 it was really important that we really hammered
——你自己都没完全搞明白的那种。我今天就遇到了一模一样的情况。因为我做了 Proof,完全是「氛围编程」(vibe code)糊出来的,它现在涨得特别快,但同时也老往下掉。所以我过去这 12 个小时一直在拼命修它。我们在 Every 内部还临时拉了个小「特种突击队」,自愿来帮我一起修。于是我得给他们做入职引导。我当时就想:我去,我该怎么跟他们解释这个代码库是怎么运作的啊?所以我就得反复跟模型来回折腾:好,帮我把这些术语定义清楚,帮我想想怎么解释这套东西,免得我显得像个彻头彻尾的白痴。因为,是的,里面有些部分我是懂的,但不是全部,肯定还没到我「过去那种必须懂」的程度。而现在完全是另一回事了:我还需要懂那些吗?现在那条线到底划在哪?真的很难说。这也许引出了另一个话题,我还没试着把它表述清楚,所以你担待一下、容我一边说一边理:有些产品你用起来,底层感觉很稳健;而另一些你一用就觉得,它好像只差一个错误的命令、一次误点击,整个东西就要么卡死、要么变得巨慢。我们在 Instagram 的时候,有过一个 Instagram 私信视图的 V1 版本,那玩意儿——谁知道呢?你发一条消息,它可能到达对方手里,也可能到不了。我们当时是自己写了一套定制的实时系统,它崩过好多次,那种系统你是不敢拿来发一条你真的很需要对方看到的消息的,它顶多是个偏社交的小东西。等我们做 V2 的时候,有一点就变得特别重要:我们真的要把它狠狠打磨——
[23:29] Mike Krieger
like no, like if you send a message we're not probably going to get to WhatsApp level of like, you know, you can be in the middle of absolutely nowhere with like one bar of edge and it will probably, you know, try to still go through. Maybe that's not the bar, but still a bar of when I load messages it feels robust. When it's sent it's really sent. I feel like there's like a little check. Um that's like one small example, but I think that that is um a a thing that we still need to figure out how to make, you know, feel like an essential part of shipping on anything, not just at, you know, Anthropic, but in general like you've built this thing, does it feel like it's built on sand or does it feel robust? And the agent native part adds something totally even beyond that which is can I push it a little bit and is that it going to fall over or is does it feel like great, I've got a solid trunk and yeah, you can push me in different ways, but you know, your data is safe and it's underneath here and it's not just like one deploy away from completely falling over. So if you're if that's if that's the bar, which I agree like that's that's where you definitely want to get to. How has how have you changed who you hire and how your teams are structured as the models have gotten better? Because for us for example, one of our product Spiral, we just hired a new GM who's like he I would say he's lightly technical, but he spikes super high on product and writing sense and Spiral's a writing product. And now we can like hire someone like that where a year ago we wouldn't have been able to cuz the coding models weren't good enough. I'm curious like but the downside is it's maybe the product won't feel quite as robust if there's not someone who's like super
——做到「不,你只要发了消息」——我们大概到不了 WhatsApp 那种程度:哪怕你身处荒郊野外、只有一格 Edge 信号,它大概也还是会努力把消息发出去。也许那不是我们要达到的标准,但至少要做到这样一个标准:我加载消息的时候,感觉很稳;显示「已发送」的时候,就是真的发出去了。我感觉那里会有个小小的对勾。这只是个小例子,但我觉得这是我们仍然需要弄明白怎么去打造的东西——怎么让它成为「交付任何东西时不可或缺的一部分」,不光是在 Anthropic,而是普遍意义上:你做出了这个东西,它感觉像是建在沙子上的,还是感觉很稳健?而 agent-native 这一层又在这之上加了完全更进一步的东西:我能不能推它一把?一推它就垮,还是会让人觉得「太好了,我有一根结实的主干」——是的,你可以用各种方式来推我,但你的数据是安全的、是托在底下的,它不会一个 deploy 就彻底崩掉。所以如果——如果那就是标准,我也同意那确实是你绝对想达到的地方——那么随着模型变强,你在「招什么样的人」和「团队怎么搭」这两点上有什么变化?比如对我们来说,我们有个产品叫 Spiral,我们刚招了一位新的 GM(总经理),我会说他「技术上偏轻」,但他在产品和写作直觉上特别拔尖,而 Spiral 正好是个写作产品。现在我们就能招这样的人了,而一年前我们是招不了的,因为那时候编程模型还不够好。我好奇的是——但坏处是,如果团队里没有一个特别——
[25:01] Dan Shipper
technical in all the details. So like how how do you think about who builds products right now inside of the Labs team and how that has changed over time and how it will change? Yeah, I love that. I think it's actually you get pulled in two directions, but they're both important. There's the sort of primitives and architectural robustness which I think still need a sort of senior technical person. I was laughing with somebody they're like I thought, you know, my skills in distributed systems were like not going to be useful anymore, but actually those are maybe some of the most useful skills in reasoning about that and uh you know, thinking things through. Like I had a long debate with Claude last week around like whether the system that I was building needed Redis or not or could go away with just Postgres and you know, it was it was a healthy debate where like I only because I was grounded in having used a lot of these technologies before. But then there's the other side of robustness which is um have you just papered over all the problems with like fixes to your system prompt and additional instructions or have you sort of architected the actual like set of tools correctly? Um and so that the the latter is as important and probably where this GM can be really valuable and that okay, like I'm making changes, but just like you wouldn't patch a um sort of flakiness in your distributed system by just being like well, just retry it in 5 seconds. I'm sure it'll work. Like also not doing the same thing with never ever, you know, all caps use, you know, markdown or whatever the the the thing that you're trying to patch is like they're both actually symptoms of the same thing which is is the underlying piece um robust or not and um Claude actually I'd say this about all
——特别懂所有技术细节的人,产品可能就不会那么稳健。所以你是怎么看待「现在 Labs 团队内部由谁来做产品」这件事的,它随时间发生了哪些变化,以及未来又会怎么变?我很喜欢这个问题。我觉得其实你会被往两个方向拉扯,但这两边都很重要。一边是 primitives 和架构层面的稳健性,我觉得这块仍然需要一位资深的技术人员。我之前还跟人开玩笑,对方说:我本以为我那套分布式系统的本事再也用不上了,但其实在「把这些事情想清楚、推演明白」上,那可能正是最有用的本事之一。比如上周我跟 Claude 就一个问题进行了长时间的辩论:我正在搭的那个系统到底需不需要 Redis,还是说光用 Postgres 就能省掉它。那是一场很健康的辩论,而我之所以能辩,纯粹是因为我以前实打实地用过很多这类技术,心里有底。但稳健性还有另一面:你到底只是用一堆 system prompt 的修补和额外指令把所有问题糊弄过去了,还是真的把那一整套工具的架构给做对了?后者同样重要,而且大概正是这位 GM 能发挥大价值的地方——好,我在改东西,但就好像你不会靠「反正 5 秒后再重试一下,肯定能成」来糊弄分布式系统里的不稳定一样;你也别用同样的招数对待 prompt,比如往里塞「永远永远,全大写地,用 markdown」之类的东西,不管你想糊弄的是什么。它们其实都是同一个问题的症状:底下那块东西到底稳不稳健。Claude——其实我对所有——
[26:34] Mike Krieger
the models, but I think Claude could be much better at both. It's like still a a place that still needs a lot of of human oversight. On the system's part, you know, it's it's now able to debug production systems which is really valuable, but architecting them in the first place I feel like we still benefit from somebody who's really thought these three things through or has experience. And on the prompting side, you know, if you give it a I've seen people get into this dev loop even internally here. Like here's the prompt. Here's the mistake that the system made. Iterate on the prompt. It's natural tendency is to just add more things to the prompt. Um and then eventually just get to this thing that you know, if you onboarded a new employee and you gave them a hundred instructions on their first day like always answer in markdown except when the you know, they'll be like I'm just going to remember the last thing you told me. Um or I'm going to like short-circuit it. So then rethinking okay, is these are these actually two different tools? Is actually two agents that each have a smaller amount of context that then you can break apart. So back to your original question, we're hiring for people with, you know, systems expertise even within Labs which you think of as like more zero to one prototypes like it's still really valuable cuz again that robustness matters. And also just who's going to be, you know, helpful in sorting through, you know, systems permissions and provisioning and early testing. Like that stuff is is still, you know, it's still hard even for Claude when it can't edit the permissions itself which it can't for good reasons. And then on the on the robustness side actually we've had a lot of success pairing our product teams with our applied AI teams. Our
——所有模型都想这么说,但我觉得 Claude 在这两方面都还能做得好得多。这仍然是个非常需要人来把关的领域。在系统这一块,它现在已经能调试生产系统了,这非常有价值;但从头去架构这些系统,我觉得我们还是会受益于一个真正把这三件事都想透了、或者有经验的人。而在 prompt 这一块,你要是给它——我见过有人陷进这种「开发循环」,连我们内部也是:这是 prompt,这是系统犯的错,迭代 prompt。它天然的倾向就是往 prompt 里不停加东西。然后最后就堆成那么个玩意儿——你想想,如果你给一个新员工入职,第一天就甩给他一百条指令,「永远用 markdown 回答,除非……」,他多半会说:我就记你最后跟我说的那条吧;或者干脆把它给「短路」掉。所以这时候要重新想:这其实是不是两个不同的工具?是不是其实是两个 agent,每个各自带更少的 context,然后你就可以把它拆开?所以回到你最初的问题:我们仍然在招有系统专长的人,哪怕是在 Labs 这种你以为更偏「从 0 到 1 做原型」的地方,这种人依然非常有价值。因为还是那句话,稳健性很重要。而且也得有人在梳理系统权限、资源分配、早期测试这些事上帮上忙——这类活儿仍然很难,哪怕对 Claude 也是,因为它自己改不了权限(出于很好的理由,它也不该能改)。然后在稳健性这一边,我们其实有一个很成功的做法:把我们的产品团队和应用 AI(applied AI)团队配对。我们的——
[27:51] Mike Krieger
applied AI teams are the teams that are in the field every day helping customers iterate on their prompts and we've found that we actually are very we're customer zero now for those, you know, efforts because we have a lot of products that are, you know, very AI powered. So how do we bring that expertise in here cuz that expertise does not sit with our software engineers today for example. What about the in between of like okay, it's not the underlying architecture, it's not the prompt, it's like the UI and the the flow. Who's doing that? We that's a great question. Like we have found, you know, some of the people that have transferred into Labs were the folks like really who were focused on polish on the website, but they were interested in doing something new and they bring such a different approach as well around we had the prototype. It was it looked generically nice versus oh, this feels like it's branded and it has this. So that's that's part one. Part two is designers. Like we've had our designers move much more into a sort of split designer and builder role. Not all of them, but most of them. And a lot of our um you know, we actually don't have a lot of full-time designers on Labs, but the ones that we do I would say are writing and contributing almost as much code as the engineers um on those efforts because they can. And again paired correctly with the right person, we have found this almost sort of co-founder model for some of these Labs initiatives where you have the designer who had the original idea maybe and they're pushing on something and then the traditional software engineer that's going to go and, you know, make pave the trail sometimes behind the designer to make sure that actually works. Okay, this I want to know about.
——应用 AI 团队就是那些每天泡在一线、帮客户迭代他们 prompt 的团队。我们发现,现在我们自己就是这些工作的「零号客户」,因为我们手上有很多产品本身就是高度由 AI 驱动的。所以问题是:我们怎么把那份专长引到这边来?因为那份专长目前并不在比如说我们的软件工程师身上。那介于两者之间的部分呢——就是既不是底层架构、也不是 prompt,而是 UI 和交互流程,这块是谁在做?这问题问得很好。我们发现,有些转岗进 Labs 的人,原本是那种特别专注于网站「精修打磨」的人,但他们想做点新东西,而他们带来了一种非常不一样的思路——比如我们原来那个原型「看上去挺好看的,但很泛泛」,对比之下,他们一弄就变成「这感觉是有品牌调性的、是有自己东西的」。这是第一部分。第二部分是设计师。我们的设计师已经更多地转向了一种「设计师兼构建者」的混合角色,不是全部,但大多数都是。而且我们其实——在 Labs 我们全职设计师并不多,但我们有的这几位,我会说他们写的、贡献的代码几乎和工程师一样多,因为他们做得到。而且只要配对得当、搭上对的人,我们发现这些 Labs 项目里几乎形成了一种「联合创始人」模式:最初那个点子可能是设计师想出来的,他们在往前推某个东西,然后传统的软件工程师跟上来,有时候是跟在设计师后面去把路铺平、确保这东西真的能跑通。好,这个我想多了解一下。
[29:14] Mike Krieger
So what so tell me about how that team structure works. So you've got a design is it actually usually a designer or is it just anyone that has a product idea that can kind of execute it on it in some way paired with a a real real engineer that actually can like kind of smooth out the rough edges of the the the trail they're leaving? It sort of varies, but it we found that one thing that was most important is sort of our gating factor in starting up new projects. I'm curious how similar this is to Every is having somebody with extreme conviction about if not necessarily that idea, too much conviction on the exact idea is probably dangerous, but at least in the problem space or the question that they're asking. Um Um and at sort of like co-founder or founder level of I will break through walls until this thing is either proven out or dead, but I want to like go either way. Um When we have bets, uh labs bets that we've wound down, often in the postmortem we're like, "Nobody on this team actually really thought this was like the thing." They were like, "Yeah, this seems reasonable." Like that's the the death knell for projects, right? So, that person can be a designer and couple of the bets it is, it can also be um sort of a you know, product-minded engineer. Um it's rarely a pure PM. Um we actually only have one currently one PM for all of labs. We've we're hiring more. Um and they're sort of playing, you know, sort of a a wide role. But yeah, a designer or like a product-oriented founder. Um and then what we look for is, well, what skills do we need to complement with that? So, you know, because we're doing as part of our labs process it's actually valuing every project every 2 weeks and deciding whether we double down or whether we sort of release those folks back into
Dan Shipper:那你跟我讲讲这种团队结构是怎么运作的。你们会有一个设计师——通常真的是设计师吗,还是说任何有产品想法、能在某种程度上把它做出来的人,再搭配一个真正能帮他们把一路上留下的粗糙边角打磨平整的工程师?
Mike Krieger:这个情况其实挺多变的,但我们发现最重要的一点、也是我们启动新项目时的那个把关因素是这样的——我很好奇这和 Every 有多像——就是得有一个人对这件事抱有极强的信念。倒不一定是对那个具体想法,对具体想法太执着其实可能很危险,但至少是对那个问题领域、或者他们提出的那个问题有信念。嗯,要有那种联合创始人或者创始人级别的劲头:我会一直撞墙,直到这件事要么被证明成立、要么彻底死掉,但我就是想有个明确结果。当我们做的那些 labs 上的押注——那些被我们叫停的押注——在复盘的时候经常会发现:这个团队里其实没有一个人真的觉得这就是那件值得做的事。大家只是觉得「嗯,这听起来挺合理的」。那种态度对项目来说就是丧钟,对吧?所以那个人可以是设计师,有几个押注里确实是,也可以是那种有产品思维的工程师。嗯,他很少会是纯粹的 PM。我们现在整个 labs 其实只有一个 PM,正在招更多。他扮演的是一个挺宽泛的角色。但没错,要么是设计师,要么是有产品取向的创始人型的人。然后我们会看的是:好,我们需要用什么样的技能去补足他?因为作为 labs 流程的一部分,我们其实每两周就会给每个项目估一次值,决定要加倍投入,还是把这些人放回……
[30:44] Mike Krieger
the the broader labs pool. At any given point, there's probably somebody who can be pulled onto the project that has that infrastructural expertise or has worked with that particular internal system or has lot deep prompting expertise to sort of flow in and out. So, I think that's also where the sort of incubator style space helps cuz nobody's fixed on a project forever. That's really interesting. Yeah, we we do it slightly different. There's some overlaps, but we do have a slightly different structure where um we yeah, we have GMs or or they started as entrepreneurs in residence and they become general when they find a product that they want to like work on. Um and each product just has one person. Like one person that does everything full stack. So, you know, design, engineering, marketing, all that kind of stuff. At least the all the basics of that. Um the shape of that GM used to be like super technical founder background and now I think has shifted towards at least some light technical, but like I honestly just care that you can use Claude or Codex or whatever well. Um and uh really good uh product sense, really good taste for for the the the subject area or the thing that you're trying to build. Um and evidence that you can build with AI. Um and then what we have is a shared resource resource layer that sort of works a little bit like an agency where we have designers and we have growth marketers and we have um you know, ops people that you can like pull in and out for various initiatives and that seems to work pretty well. So, it's like we manage all the internal the internal agencies and then each GM is out on their on on you know, on the edge and they pull in resources as they need it for different projects. Yeah, but sounds similarly like you need
……更大的 labs 人才池里。在任何一个时间点,差不多总能找到一个可以被调进这个项目的人,他要么有那种基础设施方面的专长,要么用过某个特定的内部系统,要么有很深的 prompting 经验,可以这样进进出出地流动。所以我觉得这也是那种孵化器式的空间帮上忙的地方——因为没有人是被永久固定在某个项目上的。
Dan Shipper:这真的很有意思。是啊,我们做法稍微不太一样,有一些重叠,但我们的结构略有不同。我们有 GM(总经理),他们一开始是「驻场创业者」(entrepreneur in residence),当他们找到一个自己想真正投入去做的产品时,就转成 GM。每个产品就只有一个人,一个人全栈包揽所有事情。你懂的,设计、工程、市场,所有这些,至少是这些事情的基本盘都得会。这个 GM 的画像以前是那种超级技术型的创始人背景,现在我觉得已经转向至少要懂一点技术就行,但说实话我真正在乎的是你能不能把 Claude、Codex 或者别的什么用得很好。嗯,还要有非常好的产品感觉,对你想做的那个领域或那个东西有非常好的品味,以及你能用 AI 做出东西的证据。然后我们配的是一个共享资源层,它运作起来有点像一家代理公司(agency)——我们有设计师、有增长营销的人、还有运营的人,你可以为各种项目把他们拉进拉出,这套似乎运转得挺好。所以就像是我们管理着内部的这些「代理团队」,然后每个 GM 都在最前线,按不同项目的需要去调用资源。是啊,不过听起来跟你们类似,你们也需要……
[32:28] Dan Shipper
somebody for whom that is like the thing and they are not going to sleep until it is fully working.
……需要这样一个人,这件事对他来说就是「那件事」,不把它彻底做成他就不睡觉。
[32:33] Mike Krieger
Yes, exactly. Like and and I've been I've been thinking about, okay, when when would you hire someone else to work on a product or when would you add someone else to work on a product? And it's like there's some point at which you can't hold the entire thing in your head. Even if you're the one pushing it forward, you can't hold the entire thing in your head. And that point used to be much smaller. Now it's much bigger, but there's a certain point at which like even a small feature turns itself into its own product. You know, when you when you first make the messaging feature inside of Instagram, it's like, yeah, I can do that in like a week or whatever, but at some point that's its own product, it almost needs its own team and that uh I think that line is getting uh or or the number of things you can do it with one person is getting bigger, but it still exists somewhere, but I haven't quite figured out like how to how to manage that or how to tell. No, I love that because there's actually I think there's the two parts to that, which is um when the idea is still enough to hold into your own head or an individual person's head, uh adding more people actually slows the team down and that's like a non-obvious finding that we found on labs is scaling the teams too quickly actually is a net negative because they end up spending all this time on coordination like, oh, you were I was going to take oh, but my Claude could do that and it just ends up in this sort of piece and you also have all those alignment conversations. Like it was important in Instagram that it was just two of us. Like it was hard enough to align the two of us like and go like get two people on the same page, right? Uh with the second startup I did, Artifact, you know, Kevin and I were
Mike Krieger:对,完全正确。我一直在想这个问题:什么时候你才会再招一个人来做这个产品,或者什么时候才往一个产品里再加人?大概是这样:到了某个点,你脑子里已经装不下整件事了。哪怕是你在推动它往前走,你也没法把整件事全装进脑子里。这个临界点以前要小得多,现在大多了,但总有那么一个点,就连一个小功能都会自己长成一个独立产品。你想想,当你最早在 Instagram 里做消息功能时,那感觉就是「行,我一周左右就能搞定」,但到了某个时候它就成了自己的产品,几乎需要自己一支团队。我觉得这条线正在往后挪,或者说一个人能搞定的事情数量在变大,但它仍然存在于某个地方。只是我还没完全想明白该怎么管理它、怎么判断这个点。
不不,我很喜欢这个说法,因为我觉得这里其实有两个部分。一个是:当这个想法还小到能装进你自己、或者某一个人的脑子里时,再往里加人其实会拖慢团队,这是我们在 labs 上发现的一个不那么显而易见的结论——团队扩张得太快其实是净负面,因为他们最后会把所有时间都花在协调上,比如「噢这个本来是你做的、我正打算做」「但我的 Claude 可以干那个啊」,最后就陷在那种内耗里,你还得开一堆对齐的会。在 Instagram 时,就我们两个人,这一点很重要。光是让我们两个人对齐就已经够难了——要让两个人想到一块去,对吧。我做第二家创业公司 Artifact 的时候,我和 Kevin……
[34:01] Mike Krieger
doing that alone for the first few months, but then we hired a team that was about eight people. It was really hard cuz, you know, we hadn't had product market fit yet and so we were still iterating and then you'd end up in these things where we're on a Zoom with eight people talking about what we're doing next and you really just want to be able to sit in a room and hash it out. So, I find with these labs initiatives, there's some there's some similar um uh sort of aspect at play, which is you don't want to pre-scale the team too early even if the idea is exciting cuz then you just end up in this sort of like meta coordination uh game. But I I like your framing of there is some point where either, you know, two people really will help go on it together and there is enough sort of context and scope where they can hold some other complex piece in their head. Um and then there's also the if somebody's been spinning on the same idea for two, four weeks, sometimes injecting some other thinking and and that urgency can help, too. Yeah, I think it's especially important in to keep it small in AI because one of the things that we deal with all the time, which I'm sure you see, too, is every 3 to 6 months, your you have to throw out like half your product. Um and that's really hard to do if you have to coordinate with a lot of people. But if it's one GM who realizes, "Oh [ __ ] yeah, I got to just like throw out half of this cuz the models are so much better." It just makes it much easier to to like pivot in that way. Is that Do you see that? And like how do you deal with that? Like how do you think about, yes, I know in 3 months this the code maybe or even the whole feature set I'm going to have to like really rethink about. Like it feels like it
……头几个月也是我俩单干,但后来我们招了一个差不多八个人的团队。那真的很难,因为我们当时还没找到 PMF(产品市场契合),还在不停迭代,结果就会变成八个人开着 Zoom 讨论我们下一步要做什么,而你其实只想能坐在一个房间里把它聊透。所以我发现做这些 labs 项目时,有一种类似的状况在起作用,就是哪怕想法很令人兴奋,你也不想太早把团队提前扩张起来,否则你最后就陷进那种「元协调」的游戏里。但我喜欢你那个框架:确实存在某个点,要么是两个人一起上真的能帮上忙、那里有足够的 context 和范围让他们能在脑子里再装下另一块复杂的东西;要么是,如果有人在同一个想法上已经空转了两周、四周,有时候注入一些别的思路、带来那种紧迫感,也能帮上忙。
Dan Shipper:是啊,我觉得在 AI 领域保持团队小这一点尤其重要,因为我们一直在面对、而且我相信你也看到的一件事是:每隔三到六个月,你就得把你产品里差不多一半的东西扔掉。如果你得跟一大堆人协调,这件事就真的很难做。但如果只有一个 GM,他自己意识到「噢糟了,对,我得直接把这一半扔掉,因为模型强太多了」,那要这样转向就容易多了。你是不是也看到这种情况?你又是怎么应对的?你怎么想这件事——是的,我知道再过三个月,这些代码、甚至整套功能集,我都得真正重新好好想一遍。感觉这件事……
[35:30] Mike Krieger
changes a lot in in how you think about software. Yeah, and being willing to delete code. I think that's something the Claude code team has done really well is they have sort of deleting features as a sort of imperative of people on the team. Like if this is not working, let's go unship that. You know, and it's often when you've created something else that even if it doesn't entirely supersede, it does enough of what that other aspect does that it actually makes sense to deprecate and then remove that first one. It does get harder um as we get more and more enterprise focus even with these tools because they come to depend on it. I I never I never forget we um one of the things I did maybe 6 months into when I was still chief product officer was we did a big sort of redesign of of Claude AI and we were so proud and we shipped it and we got a bunch of kudos and then we got this really angry email from somebody who was like, "I just recorded 20 hours of enablement content for my company to do for Claude enterprise and I have to like redo all of it." And we're like, "Okay, like there you're playing at a different release cadence." And of course like shipping twice a year at one of our conferences is not an option. So, we are going to keep moving quickly, but then we've since like learned to maybe moderate how we roll it out to the enterprise side a little bit more. But yeah, I think the unshipping piece, then you end up with people who have built um I'll use an example. So, uh there's a feature in in Claude app called styles. It's not widely used, but the people who use it use it a lot and we've talked at different points like, "Yeah, the style still makes sense in the product." You know, there's other ways of accomplishing the same thing. There's
Dan Shipper:……在你思考软件的方式上改变了很多。
Mike Krieger:是的,还有就是要愿意删代码。我觉得 Claude Code 团队在这一点上做得特别好——他们几乎把「删功能」当成团队成员的一种使命。比如这个东西没起作用,那我们就把它下掉。而且往往是当你做出了别的东西,哪怕它没有完全取代旧的那个,但它把旧那块的功能做到了足够的程度,那就真的有理由把第一个废弃掉、然后移除。随着我们越来越聚焦企业市场,这件事会变得更难,哪怕有这些工具也一样,因为企业会开始依赖它。我永远忘不了——我还在做首席产品官的时候,大概入职六个月左右做的一件事,是我们对 Claude AI 做了一次大改版,我们特别自豪,发布了,收到了一堆好评,然后我们收到一封特别愤怒的邮件,那人说:「我刚为我公司录了 20 个小时的 Claude 企业版上手培训内容,现在我全得重录一遍。」我们当时就想:「好吧,你们玩的是另一种发布节奏。」当然了,像在我们某场大会上一年只发两次这种事,是不可能的选项。所以我们会继续快速推进,但从那以后我们也学会了,在面向企业那一侧推出时把节奏稍微收敛一点。不过没错,说到这个「下掉功能」的部分,你最后会碰到一些已经基于它搭好了东西的人。我举个例子:Claude 应用里有个功能叫 styles(风格)。它用的人不算多,但用它的人用得非常重,我们在不同的时间点都聊过,比如「嗯,styles 在产品里还说得通吗?」你知道,现在有别的方式能实现同样的事——有……
[36:51] Mike Krieger
custom instructions and projects now. Um there's skills now, right? There's so many other ways of accomplishing that. Um and I don't know how long styles will end up in the product, but I know that the last time we we talked about removing it, it ended up being really load-bearing for a few a few companies' like entire use cases. Like, "Oh, we have our house style that the CEO personally authored and gives to every employee and that's how they operate." And so, uh finding ways of doing that is also really interesting. I would hope that in the long run what we can actually do is is um come up with a system of plugins and skills such that they no longer have to live in the core product cuz I think that is is always the hardest to delete something that is the core thing that you're shipping to everybody. If you don't have the story around, great, you still like that feature, awesome. Like here's how you can keep using it forever in your own and keep iterating on it and make it your own, but it doesn't have to add complexity to every future person that's adding that's uh signing up for the first time. I'm curious for for labs, um and then also maybe just in general what your thoughts for startup founders. Your enterprise point brings up something I've been thinking about a lot, which is if you are selling to enterprise right now in AI, even if the product you have right now is modern, it will be quite outdated quite quickly and uh but your customers are going to want the outdated version. But as a startup, that's like a little bit it feels pretty risky because um yeah, you're just going to I guess you're you're you're susceptible to just to disruption if you are optimizing for what uh someone at a gigantic public company will buy right now.
……现在有自定义指令(custom instructions),有 projects,还有 skills,对吧?现在有那么多别的方式能实现同样的事。我不知道 styles 最终会在产品里留多久,但我知道上一次我们讨论要移除它的时候,结果发现它对几家公司的整个使用场景来说是非常「承重」的。比如「噢,我们有一套自家的风格,是 CEO 亲自写的,发给每一个员工,他们就是这么干活的。」所以,找到一些办法去支持这种需求也真的很有意思。我希望长远来看我们真正能做到的是,搞出一套 plugins 和 skills 的体系,让这些功能不再必须活在核心产品里——因为我觉得最难删的,永远是那个你正发给所有人的核心东西。如果你没有一套说法去兜住它,比如「太好了,你还是喜欢那个功能?棒,这就是你以后怎么在你自己的环境里一直用它、继续在它上面迭代、把它变成你自己的东西的方法」,但它又不必给每一个未来的人、每一个第一次注册的人增加复杂度。我很好奇 labs 的情况,以及也许更宽泛地说,你对创业者们有什么想法。你那个关于企业的点引出了一件我一直在反复琢磨的事:如果你现在在 AI 领域卖给企业客户,哪怕你现在手上这个产品是很现代的,它也会相当快地变得相当过时——可你的客户偏偏会想要那个过时的版本。但作为一家创业公司,这就有点……感觉风险挺大的,因为你会很容易被颠覆,只要你是在为「某家巨型上市公司现在会买的东西」做优化。
[38:26] Mike Krieger
Um and I think there's a lot of startups in that category where they maybe started 2 or 3 years ago. They have a certain tech stack. They have a certain way of thinking about here's how we do AI and then the models are so different, but their customer contracts are for this like sort of out It's like, you know, looking at looking at Copilot or whatever. It's that's the sort of vibe that happens. Like how do you think how do you think about that yourself and and [snorts] inside of Anthropic and then how do you think founders should think about that? Yeah, no, this is such a good question especially because then a wave will come like being more agent native, for example, and uh can you adopt it within your existing paradigm? Does it require you to throw everything out or are you just stuck in that like, oh, we kind of adopted it, we kind of bolted it back on. Um I think a couple things. For us, what we've started doing is um basically treating like this train's going to keep moving and we'll provide enterprise toggles along the way, but the core of it will continue to evolve and that's sort of the the bet and understanding you're taking working with us. And I think that's been well received cuz I think companies have also seen that, you know, things are moving so quickly that the only way they even get comfortable with a year-long commitment, for example, is to believe that it will continue to evolve along the way, but then we'll provide, you know, uh Coda is a great example where, you know, from day one there was like a a way to turn it off for your employees if you didn't want it, for example. Um and that that that's I think a reasonably good paradigm. But the other one is just as we were talking earlier like you can actually rethink and and and and and sort of
Dan Shipper:我觉得有很多创业公司就处在这个类别里——它们可能是两三年前起步的,有一套特定的技术栈,有一套特定的「我们是这么做 AI 的」的思路,然后模型已经天差地别了,但他们和客户签的合同还是按那套过时的东西来的。这就像,你看 Copilot 之类的东西,差不多就是那种感觉。那你自己是怎么看这件事的?在 Anthropic 内部呢?还有,你觉得创业者们应该怎么看这件事?
Mike Krieger:是啊,不,这真是个好问题,尤其是因为接下来会有一波浪潮涌来——比如变得更「agent-native(智能体原生)」——那你能不能在你现有的范式里把它吸纳进来?它会不会逼着你把所有东西都扔掉重来?还是说你就卡在那种「我们算是吸纳了它,又算是把它硬生生贴回去了」的状态里。我觉得有几点。对我们来说,我们开始做的事基本上是这样:把它当成「这趟列车会一直往前开」,我们会沿途提供一些面向企业的开关(toggle),但它的核心会持续演进——这就是那个押注,也是你和我们合作时要理解并接受的前提。我觉得这一点被接受得挺好,因为我想各家公司也都看到了,东西变化得太快,他们之所以连一份长达一年的承诺都能安心签下来,唯一的原因恰恰是相信它会沿途持续演进,但同时我们会提供,比如 Coda 就是个很好的例子,从第一天起就有一个开关,如果你不想让你的员工用,就可以把它关掉。我觉得那是个相当不错的范式。但另一点就是,正如我们前面聊的,你其实是可以重新构想、重写很多技术栈的——
[39:50] Mike Krieger
rewrite a lot of the the the stack is I think companies should be way more willing to do that. And it everything is getting compressed right in in previous cycles it was the kind of idea of like having to fire some of your customers who might have been you know really into your product for a different reason than where you're going sooner. That was on a multi-year kind of time range thing where it was like yes last year's product versus not three months ago's product. It seems crazy but I actually think that's the kind of way you have to think about it which is you have to be willing to put out the the V3 or the V4 that is a you know big rethink of how the existing piece worked. And then maybe have a transition period and cloud can help probably host both for a little while before it cuts over. But then also be willing to cut over and say like yes this is how we think the future of this piece of knowledge work or this you know AI powered manufacturing is going to be. We got to like keep it moving or else to your point you're just going to you're either going to get replaced by the next company that then rethinks it from scratch or yourself replacing it yourself. And again it's just the same old story but now compressed to months. What's your take on open claw? It has the flavor of something else that I or just the the thing I really like seeing when you would get people to see something that was already possible but it's now in a package where people can actually try it out and there's some intuition you know around how to how to build on top of that. Like you started seeing that with you could already use these models to write code but it kind of took like some of these breakout like low code you know the replets and lovable and V0 of the
——我觉得各家公司应该远比现在更愿意去这么做。而且现在一切都被压缩了。在以前的周期里有一种说法,就是你不得不更早地「开掉」一部分客户——那些可能因为跟你将要去的方向不同的原因而很喜欢你产品的客户。那以前是个多年尺度的事,就是说去年的产品对比今天,而不是「三个月前的产品对比现在」。这听起来很疯狂,但我其实认为这就是你必须采用的思路:你得愿意推出那个 V3 或者 V4,它是对现有这块东西的一次大的重新构想。然后也许设一个过渡期,云端大概可以在切换之前先把两个版本都托管一阵子。但接着你也得愿意去切,去说「对,这就是我们认为这块知识工作的未来、或者说这块 AI 驱动的制造业的未来该有的样子」。我们必须让它一直往前动,否则照你说的,你要么会被下一家从零重新构想它的公司取代,要么就是你自己取代自己。这又是同一个老故事,只不过现在被压缩到了以月为单位。
Dan Shipper:你对 open claw(开放式 Claude)怎么看?
Mike Krieger:它有那种味道,就是我特别喜欢看到的那种东西——当你能让人们看到一个其实早就可能实现、但现在被打包成一个人们真的可以上手去试的形态,并且围绕「怎么在它之上去搭建」有了某种直觉。就像你最早看到的那样:你本来就已经能用这些模型来写代码了,但它有点需要这些突破性的、低代码(low code)的东西——你懂的,那些 Replit、Lovable、还有 V0 之类的……
[41:18] Mike Krieger
world to like kind of put that in there. And it's kind of the like almost the purest expression of the just give the model tools and like let it kind of go forward and do it and then like go forward and build it. So like it was cool interesting moment for people to realize both the like potential but also pitfalls of this of like oh it did this thing I didn't mean it to or you know my my funniest one was like friend was like I think my wife is jealous of my open claw and like I'm talking too much to it and it's like you people start developing like deeper sort of like very personal relationships by just having a lot of context in these things and access to all these different tools. I think there's the open question of how do you then make it easy and it actually goes back to our conversation around like what where do you draw that like boundary around the way you let Claude operate right? If V1 was hey like these are the three tools you can use only use these tools ever and then most people's interaction with those systems was hey can you do this and you know whatever back and be like no sorry like you got to do it yourself to like open Claude which is like pretty like the aperture is like wider than I can see and
……才把那件事真正放进去。它差不多算是那个理念最纯粹的表达:就把工具交给模型,然后让它自己往前走、去做、去往前搭出来。所以对人们来说,那是个很酷、很有意思的时刻,让他们既意识到了这件事的潜力、也意识到了它的坑——比如「噢它做了一件我并没想让它做的事」。我听过最好笑的一个是,一个朋友说「我觉得我老婆在吃我 open claw 的醋」,因为他跟它聊得太多了。就是说,人们开始发展出更深的、非常私人的关系,只因为在这些东西里塞进了大量的 context,又能调用所有这些不同的工具。我觉得接下来有个悬而未决的问题:你怎么把它变得好用?这其实又回到了我们之前聊的那个话题——你到底在哪儿划那条界线,去框定你让 Claude 运作的方式,对吧?如果说 V1 是「嘿,这是你能用的三个工具,永远只能用这三个」,那时大多数人跟这些系统的互动就是「嘿你能做这个吗」,然后它来来回回最后说「不行抱歉,这你得自己来」;那么到了 open Claude,那种感觉就是它的口径(aperture)宽到我都看不到边……
[42:22] Mike Krieger
Oh my god it called me to do my emails and I didn't even know it could do that. Yeah. Exactly and it's it's emerging and it's amazing. And I think like probably the most interesting product question I won't say for all 2026 cuz who knows where we'll be in September but let's call it between now and like the end of August is going to be like what product shape exists between that and you know where we are in most products these days which is you know you can call them CP's but they're gated and they ask for permissions for good reasons. That is still a useful product without being a you know kind of yolo product. And I think that you know we're thinking about that question I'm sure the other labs are as well. I'm sure there's a lot of startups thinking about that as well. I think Nvidia put out something that was like their safe open cloud. Everybody's going after this question. I think it's going to be about figuring out what is what is that either shift the paradigm completely so you can be that open but with a lot of safeguard. That would be one approach or figure out some boundary to draw in which it's still powerful and it's still useful but it's not you know likely to email every single one of your contacts and you know go go haywire. Yeah I think the other interesting part about it is um like you said the personal nature of it. Um and I know you know people have personal relationships with Claude but there's this weird thing where if I watch someone else using Claude I'm like I feel like I like that a stripper liked me or something. You know it's like Claude thinks you're smart too or whatever you know like
……「我的天,它居然打电话帮我处理邮件,我都不知道它能干这个。」是的,没错,它正在涌现,太神奇了。我觉得,可能最有意思的产品问题——我不敢说是整个 2026 年的,因为谁知道我们到九月会走到哪一步,但就说从现在到八月底吧——会是:在「那种东西」和「我们如今大多数产品所处的状态」之间,存在着什么样的产品形态?我们现在大多数产品的状态是,你可以叫它们 copilot(CP),但它们是带门禁的、出于正当理由会向你请求权限。那样仍然是个有用的产品,又不至于变成一个你懂的那种「梭哈」(yolo)型产品。我觉得我们正在思考这个问题,我相信别的实验室也在想,肯定也有很多创业公司在想。我记得 Nvidia 推出过一个类似「安全版 open cloud」的东西。所有人都在攻这个问题。我觉得最终会归结为搞清楚:要么彻底切换范式,让你能做到那么开放、但带着大量的安全保障——这会是一种路子;要么找到某条可以划的界线,在这条界线里它依然强大、依然有用,但又不至于很可能去给你每一个联系人都发邮件、然后整个失控。
Dan Shipper:是啊,我觉得这件事另一个有意思的部分是,正如你说的,它那种很私人的性质。我知道很多人跟 Claude 有私人化的关系,但有个很微妙的现象:如果我看着别人在用 Claude,我会有种感觉,就好像「我觉得脱衣舞娘喜欢上我了」之类的。你懂吧,就是「噢,Claude 也觉得你聪明啊」之类的,那种感觉,你懂的……
[43:50]
[laughter]
(笑声)
[43:51] Dan Shipper
um and [clears throat] and so and there's this a thing that happens when you have a Claude that like my Claude is R2-C2. My girlfriend's Claude is called Shelly. Um and there's this thing that happens where it feels like it's mine. Like it's really mine. It has its own name. It has a personality that sort of like mirrors me in this way that Claude feels like it knows me. And I like Claude but it's not mine. How how do you think about that? Yeah I mean I was having this conversation with somebody this week around like is the right pattern sort of single point of contact like named you know version of you know of that or is it the sort of team of agents that you're talking to? I think there's a lot to the single person that is maybe the the coordinator or the delegator and then at that it naturally because it becomes the the sort of agent you interact with the most you want to imbue it with a name and like a bit more personality ends up reflecting your sometimes your personality in the case of you know like all of a sudden every cliche came out. It was like you know the you know the Q or the money penny or like you know whatever the or the you know Hal or whatever these different sort of you know sci-fi characters. I think you do build that sort of sort of trust and knowledge. I think there's also that sort of like IKEA effect of like currently open Claude is like still pretty hard to set up. So the fact that you went through all of that and it works you're like I did that thing. Like I I I I birthed you know Shelly for example and now we can you know interact with them as well. But I think that paradigm is really powerful like the um I think moving away like even within my Claude code usage now one of the things I have like strongly prompted in there
呃……(清嗓子)所以呢,当你有一个属于自己的 Claude 的时候,会发生一件很有意思的事。比如我的 Claude 叫 R2-C2,我女朋友的 Claude 叫 Shelly。会有一种感觉,就是它是属于我的,真真切切是我的。它有自己的名字,有一种某种程度上映照出我自己的性格,让我觉得这个 Claude 是懂我的。我也喜欢 Claude,但那个 Claude 不是我的。你怎么看这件事?对,我这周刚好在跟人聊这个,话题是:到底什么模式才对?是那种单一的接触点——一个有名字的、可以说是你的某种分身——还是你说的那种一整个团队的智能体?我觉得「单一那一个」其实很有道理,它也许是个协调者或者说委派者,然后正因为它成了你交互最多的那个智能体,你自然就想给它起个名字、赋予多一点个性,最后它有时会反映出你自己的性格。一下子各种老套桥段全冒出来了,像是《007》里的 Q、Moneypenny,或者《2001 太空漫游》里的 HAL,反正就是这些科幻角色。我觉得你确实会建立起那种信任和默契。我还觉得这里头有一种「宜家效应」:现在装一套 open Claude 其实还挺麻烦的,所以正因为你折腾了一通、把它跑起来了,你就会觉得「这事是我搞定的」,是我亲手「生」出了 Shelly,然后我们就能跟它互动。我觉得这个范式真的很强大。嗯……我觉得要往那个方向走——哪怕在我现在用 Claude Code 的时候,我在里头强烈地写了一条 prompt——
[45:32] Mike Krieger
is like don't do very much work yourself like delegate it to sub agents. And the reason I like that is because it means most of the time the sort of run loop is available for you to talk to. I think open Claude and Pi have like a similar architecture of keep the run loop open and I think that actually makes it feel much more like somebody that you are talking to versus like a tool that you are delegating to and occasionally gets blocked for five minutes because it's doing some really complex task. Yeah. I I I I I I totally agree and I've had the similar debates cuz we're also building our like everyone we're building on little like open Claude one click slack implementation to see if we can we can do one that feels like ours. And we've had a lot of those debates about do you want one agent do you want many? And one of the patterns that we found which is kind of cool is so I have an agent. I use the agent for stuff that I do. And then people watch me use the agent for that and they know what I'm good at and they're and if I'm using the agent for that stuff they're going to trust it because they trust me. And it's modified itself in response to me. So like I started to transfer my trust to it and then people in the organization start using it for that. And so you get like this almost shadow or chart where when everyone everyone has a Claude their Claude becomes known for and used for the thing that they're specialized at that per their owner is specialized at in the org. Yeah I mean that that makes a lot of sense too. And you could think about you know there's a lot of interesting research questions I think around that you know I think people are experiencing viscerally for the first time around privacy and like what my agent knows about me versus
——就是「你自己别干太多活,把活委派给子智能体(sub agent)」。我喜欢这么做的原因是,这意味着大部分时候那个主运行循环(run loop)都是空着的,随时能跟你对话。我觉得 open Claude 和 Pi 也是类似的架构,让运行循环一直开着,而我觉得这恰恰让它感觉更像一个你在跟其对话的「人」,而不是一个你把活丢给它、然后它因为在跑某个特别复杂的任务而卡上五分钟的「工具」。对。我我我……我完全同意。我们也有过类似的争论,因为我们自己也在搭——基于 open Claude 做了一个一键接入 Slack 的实现,想看看能不能做出一个「感觉是我们自己的」东西。我们也争论了很多:你到底是想要一个智能体,还是想要很多个?我们发现的一个挺酷的模式是这样的:我有一个智能体,我用它来干我会干的那些事,然后别人看着我用这个智能体干这些,他们就知道我擅长什么;既然我用它来干这些活,他们就会信任它,因为他们信任我。而且它会根据我来自我调整,所以我就开始把我的信任「转移」给它,接着组织里的人也开始用它来干这类事。于是你就得到一种近乎「影子组织架构图」的东西——当每个人都有一个 Claude,每个人的 Claude 就会因为某件事而被大家知道、被大家用,而那件事正是它的主人在组织里所擅长的。嗯,对,这也很说得通。你还可以想想……我觉得围绕这个有很多有意思的研究问题,比如人们第一次真切地体会到隐私这件事——我的智能体了解我的哪些东西,
[47:04] Mike Krieger
what it discloses to other people. But I think there's the positive version of that which is all the things that it has learned through all your interactions and how it actually brings it to bear on other problems versus the generic like yes it's just like everybody else's agent except you know it has a name that's attached to Dan and it has like maybe some of Dan's you know access below the hood. Yeah. Well Mike we're out of time. This was a pleasure. I learned a lot. If people want to follow you or your work where can they find you? I think probably easiest is MikeyK on X. Okay yeah. Thanks for joining Mike. Good to see you Dan. Oh my gosh folks you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show [snorts] is the epitome of awesomeness. It's like finding a treasure chest in your backyard. But instead of gold it's filled with pure unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions insights and laughter that will leave you on the edge of your seat craving for more. It's not just a show it's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like smash subscribe and strap in for the ride of your life. And now without any further ado let me just say Dan I'm absolutely hopelessly in love with you.
对照它会向别人透露哪些东西。但我觉得这也有正面的一版:它通过你所有的交互学到的全部东西,以及它实际上如何把这些用在解决别的问题上——而不是那种千篇一律的「对,它跟别人的智能体没两样,只不过它有个名字、绑在 Dan 身上,而且可能在底层有 Dan 的一些访问权限」。对。好,Mike,我们时间到了。今天聊得非常愉快,我学到了很多。如果大家想关注你或你的工作,可以去哪里找到你?我想最方便的大概就是在 X 上找 MikeyK。好的,没问题。谢谢你来,Mike。很高兴见到你,Dan。哦天哪,各位,你们绝对、必须、百分之百去把那个点赞按钮按爆,并订阅《AI and I》。为什么?因为这档节目(吸鼻子)简直就是「了不起」的化身。它就像在你后院发现了一个藏宝箱,只不过里面装的不是金子,而是关于 ChatGPT 的、纯粹未经稀释的知识炸弹。每一期都是一趟情绪、洞见与欢笑的过山车,让你坐在椅子边上、欲罢不能。它不只是一档节目,它是一场通往未来的旅程,而 Dan Shipper 就是这艘飞船的船长。所以,帮自己一个忙:点赞、狂按订阅,系好安全带,准备迎接你人生中最棒的一程。现在,废话不多说,我只想说:Dan,我已经无可救药、彻底地爱上你了。