Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)
频道: Lenny's Podcast
视频: https://www.youtube.com/watch?v=tivaWTTVRhY
原文语言: en
统计: 共 169 轮 · Dianne 88 · Lenny 71
[0:00] Dianne
In 2023 when I started, nobody said anthropic and claude and coding in the same sentence.
2023 年我刚加入的时候,没人会把 Anthropic、Claude 和写代码放进同一句话里。
[0:06] Lenny
I want to go back to the beginning of anthropic. I remember dealing, man, these guys have no chance. OpenAI is so far ahead.
我想回到 Anthropic 刚起步的那段时间。我记得当时的感觉是:这帮人没戏了,OpenAI 领先太多了。
[0:14] Dianne
At the time, I saw people were starting to use these models not just for code autocomplete, but actually writing long form code and [music] sat an opportunity for us to train Opus 3 to be better at. That was the inflection. [music] I always think about Opus 45 a year later during winter break when everyone was home able to code.
当时我看到,大家开始用这些模型的方式已经不只是代码自动补全了,而是真的在写成篇的长代码——我们从中看到了机会,可以把 Opus 3 往这个方向训得更强。那就是那个拐点。我也常常想起一年之后的 Opus 4.5,那个寒假大家都在家里,人人都能写代码。
[0:31] Dianne
What was magical about Opus 45 is we also now not just had a model but a vehicle a great product experience like cloud code. Opus 45 wouldn't have had that moment without a product like cloud code and cloud code wouldn't have had that type of adoption accelerated without opus 45.
Opus 4.5 神奇的地方在于,我们手上不只有一个模型,还有一个载体——像 Claude Code 这样出色的产品体验。没有 Claude Code 这样的产品,Opus 4.5 不会有那个高光时刻;反过来,没有 Opus 4.5,Claude Code 的采用速度也不可能被推到那个量级。
[0:50] Lenny
I want to talk about how the product role is changing
我想聊聊产品这个岗位正在怎么变化。
[0:53] Dianne
for my team. The way to drive user value is to figure out the right user feedback. The evals, we actually have a saying on the team of evals are the new PRDs.
对我的团队来说,创造用户价值的路径,是找到对的用户反馈信号,也就是 evals。我们团队里有句话:evals 就是新的 PRD。
[1:02] Lenny
Something Gary Tan's been talking about. If you are willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.
这是 Garry Tan 一直在讲的一个观点:如果你现在愿意一年在 token 上花 10 万美元,那你过的就是 2028 年的人才会过的生活。
[1:10] Dianne
You have to sweat the tokens as much as you sweat the pixels. You have to be using the models to come up with good and great and better ideas. And there's no substitute for that. People need to be more ambitious with AI tools these days because they're just capable of so much.
你打磨 token 得跟打磨像素一样上心。你必须真的去用模型,用它逼出好的、很棒的、更好的想法——这件事没有任何替代品。现在大家用 AI 工具的野心还可以再大一点,因为它们能做的事实在太多了。
[1:25] Dianne
One thing I ask the team is let's say Claude 8 comes around. What changes in what users do? What does that mean for how you're building today?
我常问团队一个问题:假设 Claude 8 出来了,用户做的事情会变成什么样?那这对你今天该怎么做产品,又意味着什么?
[1:36] Lenny
Today my guest is Diane Penn, head of product for the AI research and labs teams at Anthropic. She joined Anthropic as the first technical product manager over three years ago, which is a lifetime [music] in AI time when the product team was just five engineers. She's helped ship every model adropic from claw 2 through fable. She's also helped incubate and launch claw code, MCP, skills, claw design, and also core capabilities like computer [music] use, tool use, and reasoning. It is always such a treat and so mind expanding to get to talk to someone who's at the very center of AI and product management. It's hard to imagine someone who has seen more of where things are going than the head of product for anthropics [music] research and labs teams. Before we get into it, don't forget to check out lenniesproass.com for a year free of the hottest and most beautifully crafted AI products in the world available exclusively to Lenny's newsletter subscribers. With that, I bring you Diane Penn.
今天的嘉宾是 Dianne Penn,Anthropic AI Research 与 Labs 团队的产品负责人。三年多前她加入 Anthropic,是公司第一位技术产品经理——按 AI 的时间尺度算,那简直是上辈子的事了——当时整个产品团队只有五个工程师。从 Claude 2 到 Fable,Anthropic 每一代模型的发布她都参与过。她还孵化并推出了 Claude Code、MCP、Skills、Claude Design,以及 computer use、tool use、推理这些核心能力。能跟一位身处 AI 与产品管理正中心的人聊天,永远既过瘾又开脑洞。很难想象还有谁,比 Anthropic Research 与 Labs 团队的产品负责人看过更多「接下来会发生什么」。进入正题之前,别忘了去看看 Lenny's Product Pass——全世界最火、做得最精致的一批 AI 产品,免费用一年,只对 Lenny's Newsletter 的订阅者开放。好,下面有请 Dianne Penn。
[2:36] Lenny
Diane, thank you so much for being here and welcome to the podcast.
Dianne,非常感谢你来,欢迎做客播客。
[2:40] Dianne
Thank you, Lenny. It's so nice to see you again.
谢谢你,Lenny。很高兴又见到你。
[2:43] Lenny
I want to go back to the beginning of Anthropic, uh, the early days. I remember when Anthropic first launched, this was, I don't know, years, the first model when it launched years ago, three years ago, something like that.
我想回到 Anthropic 最开始的那段日子,早期的时候。我记得 Anthropic 刚发布的时候——那大概是,我也记不太清,好几年前了,第一个模型发布,是几年前吧,三年前左右?
[2:55] Lenny
It was
差不多——
[2:56] Lenny
three years. I remember just like feeling that man these guys have no chance. Open AAI is so far ahead every like how what are they thinking? How is this possible? Open AI has won. It's too late. Uh things are very different now. The latest number I saw was Anthropic was making like I don't know $50 billion in ARR. That's like what companies used to go public at like very successful companies went public at 50 billion in valuation. Anthropic reportedly is making that every single year. You joined as one of the earliest PMs. There were something like five engineers when you joined. The model hasn't hadn't even [clears throat] launched when you joined. What was it like in those early days of Anthropic? What's something that might surprise people about what it was like at the beginning?
有三年了。我记得当时的感觉就是:这帮人没戏。OpenAI 领先太多了,他们到底怎么想的?这怎么可能追得上?OpenAI 已经赢了,太晚了。而现在情况完全不一样了。我看到的最新数字是,Anthropic 的 ARR 大概到了 500 亿美元。要知道,以前公司上市才是这个体量——非常成功的公司,上市估值也就 500 亿。而 Anthropic 据说是每年都能做出这个收入。你是最早的一批 PM 之一,你加入的时候公司大概只有五个工程师,模型都还没发布。Anthropic 早期是什么样子?有什么可能会让大家意外的地方?
[3:44] Dianne
I think a big part of what's made anthropic today actually has been very much the core of even the early days. So I joined in 2023 like you said we had five product engineers. There was one engineer for the entirety of our API business if you if you believe. Um and I think a big portion of it was the culture was really strong and I think this is something I emphasize for folks who are interested in the company. Um really do walk the walk of um the mission and the culture and the values. Um, and the energy was very much like a startup. And I think you're right. We were very much trying to find our identity in the early years. Like I think there's one piece around the technology, but how does that technology bring value to users, bring value to society, and what could it possibly be? And I think the early years were us exploring that in different ways. Like we did start with like cloud.ai I like another chat chat assistant and evolving into things like tool use. Um I think one of the moments where really we started to get into our groove was shipping things like Golden Gate Claude.
我觉得,把 Anthropic 造就成今天这样的很多东西,其实在最早期就已经是内核了。像你说的,我 2023 年加入,当时有五个产品工程师。整个 API 业务只有一个工程师,你敢信。我觉得很关键的一点是文化非常强——这也是我经常跟对公司感兴趣的人强调的:使命、文化、价值观,是真的照着做,不是挂在墙上的。那时候的氛围就是一家创业公司。你说得对,早期我们确实一直在找自己的定位。技术是一块,但这项技术怎么给用户创造价值、给社会创造价值,它到底能变成什么?早期那几年,我们就是在用各种方式探索这个问题。我们一开始做的是 Claude.ai,跟别人一样是个聊天助手,然后慢慢演进到 tool use 这类东西。我觉得真正让我们开始摸到手感的一个时刻,是做出了 Golden Gate Claude 这种东西。
[5:00] Dianne
I don't know if you like remember that. No.
不知道你还记不记得那个。(Lenny:没印象。)
[5:02] Dianne
Um so this this was actually up for about 24 hours or so. Uh we had just published one of our um early interpretability research in early 2024. And one of the examples was essentially you could have what's called like features of the model within the layers which uh express certain types of uh thematics. So one of the one of the themes that the researchers was able to identify was uh let's say bullet point writing. Another one was people and places. And one that really came up frequently that uh resonated was the Golden Gate Bridge. And so when you actually uh essentially dialed up that feature, Claude would obsess about the Golden Gate Bridge. So meaning in every one of its responses, it would come back and talk about the Golden Gate Bridge. So if you said like, "Give me a recipe for making spaghetti." Uh it would say, "Here is a recipe, and the orange color is just like international red that the Golden Bridge, Golden Gate Bridge looked like." Um, and so it was like really quirky and we we we very much wanted to in that situation just bring that user bring bring it to the masses and bring it to people who are starting to use claude and uh so the entire uh experience actually we spun up on our cloud.ai I website within 24 hours and that took like engineering, product, design, uh our like research teams all working together and we were really really proud of it. I think it maybe reach only 2,000 people to [laughter] be honest. Uh but it it made us feel like oh we can actually bring new user experiences, showcase our research in a way that's different and authentic to us and in a very startupy like pace. That
这个东西其实只上线了大概 24 小时。当时我们刚发布了 2024 年初的一项早期可解释性研究。其中一个例子是:模型的网络层里存在所谓的「特征(features)」,它们对应某些主题。研究员识别出来的主题里,有一个是「用要点列表写东西」,另一个是「人物和地点」,而反复出现、特别有共鸣的一个,是金门大桥。所以当你把那个特征的强度调高,Claude 就会对金门大桥着魔——意思是它每一条回答都会绕回去讲金门大桥。比如你说「给我一个做意面的菜谱」,它会说「这是菜谱,而那个橙色就跟金门大桥的国际橘一模一样」。特别有意思,所以当时我们特别想把它推给大众,推给刚开始用 Claude 的人。整个体验我们在 24 小时内就在 Claude.ai 上线了,工程、产品、设计、研究团队全部一起上,我们真的非常自豪。说实话,最后可能只触达了两千个人。但它让我们意识到:哦,原来我们真的能做出新的用户体验,能用一种不一样的、属于我们自己的方式把研究成果展示出来,而且是以创业公司那种速度。那件事——
[6:56] Dianne
to me was like one of those like maybe hidden inflection points of we were starting to find our identity that we could build products, build experiences that were different for what our competitors had seen, what was already out there. And I think that obviously labs, clog code, etc. Like we then started to identify ourselves as what we actually think the world uh how to think about AI, how to bring that closer to the public. Um but it was a very bottoms up culture. And so that entire experience was very bottoms up. I see engineers, I see uh designers donating time to work on. Um, and so I I like to always use that as example of like what the day early days were like, but the culture and and and the values have very much I think stayed the same since those early days.
——在我看来,算是一个隐形的拐点:我们开始找到自己的身份认同,意识到我们是能做产品的,能做出跟竞品、跟市面上已有的东西都不一样的体验。后来的 Labs、Claude Code 这些,显然也是从这条线下来的。从那以后,我们开始把自己定义成:我们真正认为世界该怎么理解 AI、怎么把 AI 拉得离大众更近。而且那是一种非常自下而上的文化,整件事就是彻底自下而上做出来的。我看到工程师、设计师主动贡献自己的时间来做这个。所以我特别喜欢拿它举例说明早期是什么样子——但文化和价值观,我觉得从那时到现在基本没变。
[7:46] Lenny
This episode is brought to you by our season's presenting sponsor work OS. What do OpenAI, Anthropic, Cursor, Versell, Replet, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by work OS. If you're building a product for the enterprise, you've felt the pain of integrating single signon, skim, arback, audit, logs, and other [music] features required by large companies. Work OS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SAS. Literally, every startup that I'm an investor in that starts to expand upmarket ends up working with work OS. And that's because they are the best. Whether you are a seedstage startup trying to land your first enterprise customer or a unicorn expanding globally, work OS is the fastest path to becoming enterprise ready and unblocking [music] growth. It's essentially Stripe for enterprise features. Visit workos.com to get started or just hit up their Slack where they have actual engineers waiting to answer your questions. Workos allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience. Go to works.com to make your app enterprise ready today.
本期节目由本季的冠名赞助商 WorkOS 带来。OpenAI、Anthropic、Cursor、Vercel、Replit、Sierra、Clay,以及其他几百家跑出来的公司,有什么共同点?它们背后跑的都是 WorkOS。如果你在做面向企业的产品,一定体会过接入单点登录(SSO)、SCIM、RBAC、审计日志这些大客户必备功能的痛苦。WorkOS 把这些卡单的环节变成了可以直接接入的 API,它的现代化开发者平台就是专门为 B2B SaaS 打造的。说真的,我投的创业公司里,只要开始往上走去打大客户,最后都会用上 WorkOS,原因就是它确实是最好的。不管你是刚拿种子轮、正想拿下第一个企业客户,还是已经成了独角兽要做全球扩张,WorkOS 都是最快达到「企业级就绪」、把增长堵点打通的那条路。它本质上就是企业级功能界的 Stripe。想开始用就去 workos.com,或者干脆去他们的 Slack,那儿有真正的工程师在等着回答你的问题。WorkOS 让你开发得更快——API 顺手、文档齐全、开发体验流畅。现在就去 workos.com,让你的产品今天就达到企业级标准。
[8:56] Lenny
What are some of the other um big inflection moments as you think about just Anthropic going from just this like lab that's trying to compete with this juggernaut of OpenAI at that point to what it is today? What are some moments that stick out of like wow that really changed things? Definitely when we were training and uh testing uh Opus 3, I think that was the moment when the company I think we were less than 200 people still at that point and it was very clear that we needed and wanted to create a frontier model and a uh that was very important in terms of like our ability to reach like users, consumers and uh to showcase our research. And we were looking for ways for also why should somebody choose Claude? And that was like a core question and that was a core question we were getting asked in the early days. And I think with Opus 3, you know, it launched I think early March 2024. But there was many many months of various teams across inference across research fine-tuning pre-training that rallied at different points and towards a common goal and uh I think everybody that was involved was like really proud. I remember uh being the PM, us uh the research leagues, myself, we were all in our um this was around December, so we were all at home in our various uh um parents' homes and seeing everybody's background of like their childhood room and everybody was working really hard uh to figure out that like what are we training the model for? Is it showing up the right way? So I think that was really powerful in terms of just building a lot of trust and a lot of our research leads have actually uh from that time are now like
还有哪些别的重大转折时刻?我是说,Anthropic 从当年那个想跟 OpenAI 这个庞然大物竞争的实验室,走到今天这个样子,有哪些瞬间让你觉得「哇,这真的改变了一切」?
Dianne:肯定要说训练和测试 Opus 3 的那段时间。我觉得那是一个节点——当时公司还不到 200 人,但大家都非常清楚:我们需要、也想要做出一个前沿模型。这件事对我们能不能触达用户、触达普通消费者,以及能不能把我们的研究展示出来,都特别重要。同时我们也一直在找答案:别人凭什么要选 Claude?这在早期是一个核心问题,也是外界不断在问我们的问题。Opus 3 是 2024 年 3 月初发布的,但在那之前有很多很多个月,推理、研究、微调、预训练各个团队在不同节点集结起来,朝着同一个目标使劲。我觉得所有参与过的人后来都特别自豪。我记得当时我是 PM,我和几位研究负责人——那大概是 12 月,我们都各自在家、在父母家里,视频里能看到每个人身后是自己儿时的房间,所有人都在拼命琢磨:我们到底要把模型训练成什么样?它现在的表现对不对?所以那段经历特别有力量,建立起了大量的信任。当年那批研究负责人,现在很多都在……
[10:49] Dianne
leading reinforcement learning leading our character work alignment work. So that that foundational trust I think also helped us work well now with any of our production models across product and research because we were working just so much in the trenches together in the early days. And then I think there were things like identifying that coding was important. Right? In 2023 when I started um nobody said anthropic and claude and coding in the same sentence. I think competitor models like GPT4 at the time was used a bit for coding but it was one of many use cases. And one thing that for example I saw was people are starting to use code uh these models not just for code not just like code autocomplete but actually writing long form code and is that an opportunity for us to train you know opus 3 to be better at and it ended up being a relatively smaller change from a training perspective but it ended up helping us differentiate in the early days uh competitively for users. and actually bring a lot of the very early cla enthusiasts and developers because we were uh providing a value that they didn't really think was possible at the time.
……带强化学习,带我们的角色(character)工作和对齐工作。所以那种打底的信任,让我们现在在产品和研究之间协作任何一代线上模型时都很顺畅——因为早期我们真的是一起在战壕里泡了那么久。
再往后还有一些时刻,比如意识到「写代码」这件事很重要。2023 年我刚来的时候,没有人会把 Anthropic、Claude 和 coding 放进同一句话里。当时像 GPT-4 这样的竞品模型确实有人拿来写代码,但那只是众多用例中的一个。而我当时观察到的一件事是:大家开始用这些模型,不只是做代码自动补全,而是真的在写成段成篇的代码。那这是不是一个机会,值得我们把 Opus 3 往这个方向训得更强?结果从训练的角度看,这个改动其实相对不大,但它在早期帮我们在竞争中做出了差异化,也真的带来了一大批最早期的 Claude 死忠用户和开发者——因为我们提供了一种他们当时觉得根本不可能实现的价值。
[12:10] Lenny
It's so interesting you talk about Opus 3 like that's so long ago and just like it's hard to think that was a big inflection and so this is really interesting to hear that that was internally a big milestone. It almost feels like this confidence you all built that wow we could really ship a frontier model which is now today so not great if you compare it to what we've got today. What I always think about is Opus 45 which was and interestingly like a year later also during winter break when everyone was home able to code. Uh was that another big milestone?
很有意思,你居然在讲 Opus 3——感觉那已经是好久以前的事了,很难想到它当时是个大转折点,所以听到它在公司内部是个重要里程碑,特别有意思。那几乎像是你们建立起了一种信心:哇,我们真的能做出前沿模型——虽然拿今天的东西一比,那个模型已经不怎么样了。我一直在想的是 Opus 4.5,有意思的是那也差不多是一年之后,同样是冬假期间,大家都在家里、都有空写代码。那是不是又一个大的里程碑?
[12:38] Dianne
Yeah. Um Opus 45 was definitely another large moment. I think what was magic about magical about Opus 45 is we also now not just had a model but a vehicle which is like a great product experience like cloud code. Um one thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models. And I think you know we felt the magic of cloud code for very for for uh for many months before that. Uh but the fact that the model essentially got to a level of intelligence where at a very broad level users can experience both frontier intelligence in new use cases allow it to run things end to end in an agent manner. I think that was the inflection. It was actually both. I I think Opus 45 wouldn't have had that moment without a product like Cloud Code and Cloud Code I think wouldn't have had that type of adoption accelerated without Opus45.
是的。Opus 4.5 绝对是另一个大时刻。我觉得 Opus 4.5 神奇的地方在于,我们那时不只有一个模型,还有了一个载体——一个非常好的产品体验,也就是 Claude Code。我们团队里常说一句话:你得先有前沿的产品,才能撑得起前沿的模型,才能让人真正感受到前沿模型的魔力。其实在那之前好几个月,我们内部就已经感受到 Claude Code 的魔力了。但真正的转折点是,模型的智能水平终于到了一个程度:让非常广泛的用户既能在新场景里体验到前沿智能,又能让它以 agent 的方式端到端地把事情跑完。所以这其实是两件事共同作用的结果:我觉得没有 Claude Code 这样的产品,Opus 4.5 不会有那个高光时刻;反过来,没有 Opus 4.5,Claude Code 的采用速度也不会被推到那个程度。
[13:50] Lenny
So kind of speaking on on this on this thread uh Daario interestingly if you look back at all his predictions he's just like okay coding is going to be solved it 100% in like a year something like that. He kept talking about how we're going to do code like AI is going to do all our code. And I remember everyone uh being like, "There's no way. This is way too complicated. How is how is AI ever going to get really good at this very complex thing that humans do? No, this is going to be humans for a long time." He was completely right. Something else that he talks a lot about is this exponential that we're now that we're on. That's the way he describes it now. We're like, we're on the exponential curve. I remember not long ago we were new models were being released and everybody was like, "Okay, we're done. There's no more upside. It's plateauing. It's over. There's no more room to grow." Uh, and now it's like the opposite. Now we're inside, like if you think about the curve of the exponential. We're like inside of the exponential now, which by definition means every improvement is a massive jump because we're like on that hockey stick part. What's it like just being on the inside of this crazy historic moment when AI is improving so fast, so much is being unlocked? uh what is it like and how should people prepare for the coming acceleration of more and more improvement from AI? One thing I like to say on the team is most of us weren't like actively working yet when the internet transitioned from this novelty to something that everyone can use and it feels like that's just taking humans uh I think analogies are helpful and so
顺着这条线说,有意思的是 Dario——你现在回头看他所有的预测,他就是那种「好,写代码这件事一年之内会被 100% 解决」的说法。他一直在讲 AI 会把我们的代码全写了。我记得当时所有人都说:不可能,这太复杂了,AI 怎么可能在人类做的这么复杂的事情上做到特别好?不行,这事还得靠人做很久。结果他完全说对了。
他现在还常讲的另一件事,是我们正身处的这条指数曲线——他现在就是这么形容的:我们在指数曲线上。我记得不久之前,新模型一个接一个发布,所有人都说:好了,到头了,没有上升空间了,已经进入平台期了,结束了,涨不动了。而现在完全反过来。现在我们是在——你想象一下那条指数曲线——我们现在是在指数曲线的内部,按定义讲,这意味着每一次进步都是一次巨大的跃迁,因为我们正处在那个曲棍球杆往上翘的那一段。
身处这个疯狂的历史时刻内部是什么感觉?AI 进步这么快,这么多东西被解锁。那是种什么体验?另外,面对接下来 AI 会越来越快的进步,大家该怎么准备?
Dianne:我在团队里常说的一句话是:互联网从一个新奇玩意变成人人都能用的东西的那个阶段,我们大多数人都还没开始工作。现在的感觉就像那样,只是这次它正落到人类身上……我觉得类比是有帮助的,所以……
[15:23] Dianne
like the analogy of that is I think a couple of things um number one is adaptability becomes very important um I Think we we have evals. We have you know on the safety side safety testing red teaming on the capabilities and product side new prototypes products like cloud code tag and others but it's very hard to predict the exact moment or the exact model and so the adaptability of when you're faced with new information how do you then make better decisions versus keeping the same plan. And so like that agility is really important. I think another piece is with that how do you actually be thinking very first principles and reason through what's next? What's the so what? How do we invest in new products? How do we invest in explaining the differences to users? So a lot of the a lot of the experiences I think of being in that exponential is that pace understanding how you operate and make better decisions and then applying that first principles thinking to then do something that maybe we pull up a plan that uh we would were expecting a few months from now but now the model can actually do uh and work on and actually bring that to user. So this is things like co-work skills tag, you know, as the it's a very positive self-reinforcing loop. And I I I think a big part of it also is just having the like trust in each other like making sure we have like we're we're thinking through the right decision making. We're bringing folks along. Some teams might see the exponential feel it faster than others. So how do we kind of have the grace to bring the organization, the growing organization and company along on that?
……这个类比能引出几件事。第一,适应力变得特别重要。我们有各种 evals,安全那边有安全测试、红队演练;能力和产品这边有新的原型和产品,比如 Claude Code、Claude Tag 等等。但你很难预测究竟是哪一个时刻、哪一个模型会带来突破。所以适应力的关键在于:当新信息摆到你面前时,你是继续照原计划走,还是据此做出更好的决策。这种敏捷特别重要。
第二点是,在这个基础上你要非常第一性原理地思考,推演接下来会发生什么。那又怎样?这意味着什么?我们要怎么投入去做新产品?要怎么投入去把这些差异讲清楚给用户?所以身处指数曲线里的很多体验,就是那个节奏——搞清楚你怎么运作、怎么做出更好的决策,然后用第一性原理去做事,比如把一个本来预计几个月之后才做的计划提前拿出来,因为模型现在真的能做到了,那就赶紧做出来交到用户手里。像 Cowork、Skills、Tag 就是这么来的。这是一个非常正向的自我强化循环。
我觉得还有很大一部分在于彼此之间的信任——确保我们把决策想清楚,确保把人带上车。有些团队会比其他团队更早、更强烈地感受到这条指数曲线,那我们怎么有那份耐心和体面,把这个不断变大的组织和公司一起带上路?
[17:22] Lenny
So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on is being adaptable as things emerge. Uh to your point, the product itself has to stay up to has to catch up to what is possible. To your point again, just like it can do so much, but people may not understand how to do it and may not be able to do it. So the product making it easy and even just like telling you here's something you could do feels like an important part. Is that roughly what you're describing?
所以我听下来是:几乎每一次模型发布,你们自己也不知道它会带来哪些新的可能。因此最该抓的是——随着新能力冒出来,你要足够能适应。另外按你说的,产品本身必须跟上模型已经能做到的事。还有一点,模型能做的事很多,但用户可能不知道怎么用、也用不起来,所以产品要把它变简单,甚至主动告诉你「你可以这么用」,这似乎是很重要的一环。我这么理解大致对吗?
[17:52] Dianne
I I think so. I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling loss in in the lens of um as you add in more compute and data what's called loss aka the loss from next token prediction uh goes down. And so it's a very smooth linear curve of like the models get more intelligent as you scale them up. What's actually also interesting uh in that paper is there are these like very uh different emerging capability graphs. And so for example uh as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact moment like you need the ebells to be able to assess that has actually always been a part of uh how this technology works and also what makes like things like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know M that's so interesting that you may have developed this like AI brain that uh can do something you're not even aware of and so part of the job is just uncovering wow we just got really good at this thing what can we do with that
我觉得是的。最早那批 scaling law 的论文里有几张图很有意思。大家比较熟悉的 scaling law 是这个视角:算力和数据加得越多,所谓的 loss(也就是下一个 token 预测的损失)就越低。它是一条非常平滑的线性曲线,意思是模型规模越大就越聪明。但那篇论文里还有一点也很有意思:里面有一批「涌现能力」的图长得完全不一样。比如说,随着你加入更多数据、用更多算力去训练模型,你会看到能力其实是以不连续的方式跳上去的。模型可能从「1+1 都算不对」,一下子变成「稳定能算对」。所以这种涌现能力、这种可预测性上的特点意味着:并不是每个人都知道究竟是哪一刻发生的——你得有 evals 才评估得出来。这件事从来都是这项技术运作方式的一部分,也正是它让安全变得更难:因为如果你没有那个 eval、没有那套测试系统,这些跃迁可能已经发生了,而你并不知道。
Lenny:嗯,这太有意思了——你可能已经造出了一个 AI 大脑,它能做的某些事情你自己都还没意识到。所以工作的一部分就是去发掘:哇,我们在这件事上突然变得很强了,那能拿它来干什么?
[19:27] Dianne
I think there's like product overhang and user overhang like to to maybe put it in our um PM language even on today's models and I think there's like a lot that uh we could be exploring on like our current opuses and definitely with like Fable for example temple and that that discovery is actually another part of what's been in the early days of anthropics DNA and I think is also continuing to be a big part of how we operate in product in labs and and across research. This makes me think about something Gary Tan's been talking about uh president of YC. I don't know what his title is. uh he's he had this interesting point that if you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live because by then it'll be really cheap. Everyone can work this way. But if you there's this alpha opportunity right now to just live in the future, go crazy on token spend. Uh and so there's a big opportunity for people to learn what the future's like and also just build much faster. Thoughts on this idea of and the value of token maxing, let's call it.
用我们 PM 的话来说,我觉得存在「产品侧的能力过剩」和「用户侧的能力过剩」(product overhang / user overhang)。即便是在今天已有的模型上——我觉得在现在这几个 Opus 上,尤其是像 Fable 这样的新模型上——还有很多东西是我们可以去挖的。而这种「发掘」本身,其实也是 Anthropic 早期 DNA 的一部分,现在仍然是我们在产品、在 Labs、在整个研究上运作方式的重要一环。
Lenny:这让我想到 Gary Tan 一直在讲的一件事——他是 YC 的总裁,我不太确定他现在的头衔具体是什么。他有个很有意思的观点:如果你现在就愿意一年在 token 上花十万美元,那你过的就是 2028 年的人才会过的生活,因为到那时这些会变得非常便宜,人人都能这么干。但现在,你有一个「提前住进未来」的超额收益机会——就是疯狂地烧 token。所以现在人们有个很大的机会:既能提前体会未来是什么样,也能让自己做东西快得多。你怎么看这个「token maxing」(把 token 花到极致)的价值?
[20:37] Dianne
Yeah, I think I I take more of like a almost product lens. It's almost like token spin is more the input and really the output is what you described of experimentation and I think if we were orienting like goals around experimentation. I feel like that that might be the better framing of the outcomes and therefore there might be different ways of achieving that outcome. I will say internally some of the most creative thinkers, the best like prototypers do spend a lot of time with Claude with every new version of a research model that we have. And so there is something around you have to be like using the models to then come up with good then great than better ideas and there's no substitute for that. um it's very hard to come up with a perfect strategy without touching the technology when it's moving this quickly. At the same time, I think there's other things that we could be doing like so one thing that we um do a lot is actually working in public at internally within anthropic. And so in the early days when we had less product surfaces, there was a slack channel where everyone almost the entire company was testing early versions of Claude and trying different use cases. Like people were not calling them use cases, but you might be asking it to edit an essay or uh to come up with the right way to send this email. Like they were all different use cases, but we all worked in public.
我更愿意从产品的视角看这件事。token 花销更像是输入,真正的产出其实是你说的那个「实验」。我觉得如果我们把目标设在实验上,可能是对结果更好的一种表述方式,而达成这个结果的路径也可能不止一条。
不过我得说,公司内部那些最有创造力的人、最会做原型的人,确实会花大量时间跟 Claude 待在一起——每一版新的研究模型出来他们都在用。所以确实存在这么一点:你必须真的在用这些模型,才能想出好点子,然后是很棒的点子、更好的点子,这一点没有替代方案。在技术变化这么快的时候,不上手碰这项技术,是很难想出一个完美战略的。
同时我觉得还有别的事情可以做。比如我们在 Anthropic 内部很常做的一件事,是「公开地工作」。早期我们的产品界面还很少的时候,有一个 Slack 频道,几乎全公司的人都在里面测试 Claude 的早期版本、尝试各种用法。大家当时不会管那叫「用例」,但你可能是在让它改一篇文章,或者让它帮你想怎么把这封邮件写得更得体——其实都是不同的用例,只不过我们都是在公开地做。
[22:14] Dianne
And then what you would see magically is different users or different different folks on the team coming up with an idea and then other people trying different variations of that idea and then within maybe 10 or so requests there was something magical or potentially a new use case that emerges. And I think there's a lot in not just individuals figuring out by themselves how to use this technology. I think we could be doing more to actually bring like that communal discovery when we do experimentation. Like experimentation is not always necessarily a individual sport.
然后你会神奇地看到:不同的用户、团队里不同的人抛出一个想法,接着别人会去试这个想法的各种变体,大概十来次请求之后,就会冒出某个很神奇的东西,或者一个全新的用例。所以我觉得,这里面的关键不只是每个人自己摸索怎么用这项技术。我们其实可以做得更多,在做实验的时候把那种「群体性的发现」带进来。实验并不一定总是一项个人运动。
[22:55] Lenny
It's so interesting. Yeah. This idea that we're just we're not sure what this is capable of or what we could do with it and it takes all this poking around and people trying things, hearing what other people are trying to figure out what's possible. such an interesting I don't know technology or just like okay here's what oh I figured out it could do this thing what are you gonna do with that
太有意思了。是啊,就是这种感觉——我们并不确定它到底能做什么、我们能拿它做什么,需要大量的四处试探,需要人们动手尝试,还要听别人都在试什么,才能弄明白什么是可能的。这真是一种……我不知道该怎么形容,一种很特别的技术,就像:喏,我发现它能干这个了,那你打算拿它干嘛?
[23:13] Dianne
I think at a broad theme we know right we know that the models to write great essays or you can write long form writing but individual pain points of what can you actually solve with that and bring it to like a user level that people can use um I think is something that is more exploration or experimentation uh based
我觉得在大方向上我们是知道的——我们知道模型能写出很好的文章,能做长文本写作。但具体到有哪些真实的痛点是它能解决的、怎么把它做到用户真正能用起来的程度,我觉得这部分更多要靠探索和实验。
[23:33] Lenny
following this thread you uh you oversee product for the labs team which uh is extremely cool. We've had Ben man on the podcast, Mike Griger who whom both work on labs now talk about labs. What is labs? What's come out of labs? Many people have heard of these things and how do they work that enables them to create such innovative ideas outside of even the core anthropic product team. The thesis of labs in many ways is identifying and pulling the thread on the thread of discontinuous large bets that might not be in the core road map and figuring out is there a there there and also what is the 10x 100x a thousandx of the there there and so for example uh things like cloud code um I think
顺着这条线聊——你负责 Labs 团队的产品,这特别酷。我们请 Ben Mann 和 Mike Krieger 上过播客,他们现在都在 Labs,也聊过 Labs。Labs 到底是什么?从 Labs 出来过哪些东西?这些产品很多人都听说过,但 Labs 是怎么运作的,才让它能在 Anthropic 核心产品团队之外做出这么多创新?
Dianne:Labs 的核心命题,很大程度上就是去识别并深挖那些「不连续的大赌注」——这些赌注可能并不在核心路线图上——然后判断它到底成不成立,以及如果成立,它的 10 倍、100 倍、1000 倍形态会是什么样。举个例子,像 Claude Code 这样的东西,我觉得……
[24:26] Lenny
I've heard of [laughter] uh things like cloud code uh things like uh skills and most recently cloud design MCP the thing that we really try to emphasize within the teams is especially right now there are so many things that could be built what does it mean then to have a discontinuous bet and I think one approach that we're taking this year is you can be very strongly held opinion about the theme or the area and then more weekly held about the exact prototype. And so like there is a culture of experimentation. Um there's a lot of the bottoms up like engineers on the team are very selfable um self-driven to test out different ideas and sometimes uh we have a thesis and it might not work yet and so we then might revisit it in one to two model generations. And so this idea of like these prototypes that actually end up just helping us learn like that's also valuable even if it doesn't lead to something immediately shipping. And so I think that allows the incubation and like the charter of labs to really accelerate and see around corners more broadly for anthropic. It's so funny to think about a labs within an anthropic which is already so innovative and and creative and just you know shipping like crazy that there's value to still creating a labs team within anthropic.
这个我听说过(笑)。
Dianne:像 Claude Code、像 Skills,还有最近的 Claude Design、MCP。我们在团队里特别想强调的一点是:尤其在当下,能做的东西太多了,那么「一个不连续的赌注」到底意味着什么?我们今年采取的一个做法是:对主题、对方向可以持有非常坚定的观点,但对具体做成什么原型则保持相对松的态度。所以团队里有一种做实验的文化。有大量自下而上的东西——团队里的工程师非常自驱,会自己去试各种想法。有时候我们有一个判断,但当下还跑不通,那我们可能会等一到两代模型之后再回头看。所以「原型本身就能让我们学到东西」这件事是有价值的,哪怕它没有立刻变成能发布的产品。我觉得正是这一点,让 Labs 的孵化和使命能真正加速,也能替 Anthropic 更广泛地提前看到拐角后面的东西。
Lenny:挺好笑的——在 Anthropic 这样一家本身已经很创新、很有创造力、疯狂发版的公司里,居然还需要再设一个 Labs 团队,而且这么做还确实有价值。
[25:56] Lenny
What enables labs to work as well as it has because you listed all these products and it's let's like what else has anthropic shipped it like feels like all the biggest wins almost. I'm sure there are many that I'm not thinking about right now. What's what's kind of core to creating a successful labs or within within a larger company? I think that team culture like similar to broadly at anthropic I think that team culture is very valuable. I think Ben sets an uh incredible vision and pushes people to think about the 10x 100x of the idea and you know our the teams the pods within labs is small. Sometimes these ideas start with one engineer, right? And I think uh sometimes when there's almost really large teams pursuing very ambiguous large ideas, you end up actually being slowed down because of that. Um so I think it's culture. I think you know we actually also select for folks who actually want to do that zero to one experimentation and it's not easy. There's a lot of bets that we end up turning down or turning off. Um and maybe you know we revisit them in the future. Uh but that's hard. That's hard when you pour your heart and soul.
是什么让 Labs 能做得这么好?因为你刚才列的那些产品——你会想,Anthropic 还发过什么?感觉几乎所有最大的胜仗都出自这儿。当然肯定还有很多我一时没想起来的。在一家更大的公司内部,要办好一个 Labs,核心是什么?
Dianne:我觉得团队文化很关键,这跟 Anthropic 整体是一样的。Ben 定了一个非常了不起的愿景,逼着大家去想一个想法的 10 倍、100 倍会是什么样。而且我们 Labs 里的小组都很小,有些想法一开始就只有一个工程师在做。我觉得有时候一个特别大的团队去追一个非常模糊的大想法,反而会因为人多而变慢。所以我觉得首先是文化。另外我们在招人上也确实是在筛那些真正想做 0 到 1 实验的人,而这并不容易。我们最后会砍掉、关掉很多赌注,也许未来某天会再捡起来。但这很难——当你倾注了全部心血……
[27:18] Dianne
You're acting as a founder for a bet and it's not working yet. Um so I think it's like that type selecting for that type of personality folks who are really passionate and deep about the zero to one.
你其实是在以创始人的身份押一个赌注,而它暂时还没跑通。所以我觉得关键就是筛出这一类性格的人——那些对 0 到 1 真正有热情、能钻得很深的人。
[27:30] Lenny
So you lead product for the research team. You work with the researchers at anthropic. A lot of people kind of get an sense of what is research what are research what researchers do. I think a lot of people don't totally understand these very valuable people uh at all the AI labs. Uh the way I think about it and I want to help people understand help me understand just what are researchers doing all day. What I imagine is they have a hypothesis for how to improve the model. They find data, they tweak some algorithms, they check adjust how it's trained, and they test it, see how it did, keep iterating, and keep trying to find ways to improve the model. Is that roughly right slash help us understand what researchers are doing all day?
你负责研究团队的产品,跟 Anthropic 的研究员们一起工作。很多人对「研究是什么、研究员在做什么」只有一个模糊的印象。我觉得大部分人并不真正理解各家 AI 实验室里这些极其宝贵的人到底在干嘛。我自己是这么理解的,也想请你帮大家理解一下:研究员一整天到底在做什么?我想象的是,他们对怎么改进模型有个假设,然后去找数据、调一些算法、检查并调整训练方式,然后做测试、看效果如何,不断迭代,不断寻找让模型变强的办法。大致是这样吗?或者说,帮我们理解一下研究员一整天都在干什么?
[28:08] Dianne
That's really I I think that's a lot of uh maybe the the like the more day-to-day. I think one piece around uh researchers and like research organizations like at anthropic is there's also a vision of the future like more broadly. So for example things like uh I think even at the founding of the company researchers were talking about how do we get cla to you use a computer how do we get AI to like navigate a screen right so there's a lot of actually very founderlike energy is how I describe it within researchers or really bold and ambitious researchers um and we have a ton of those at at anthropic so there's one layer of vision of what this technology can go and then I think on this other side of the loop there's also now that this technology or cloud is in people's hands how do we make it better today so it's a medium and long term and a lot of energy thinking about that lens of the future and also in the immediate and short term what are the improvement areas we can make and so like I think you're describing a really good sense of how do we make iterative improvements on different versions of claude the way that like my team works with researchers is kind of being very integrated and embedded in in those loops particularly areas where there's a lot of impact on users. So this is things like vision, computer use, coding, agent coding, tool use, test time, compute, things where there's a direct user impact and then figuring out what are the ways to uh bring the user feedback and ground it in a level that is understandable for user uh for researchers and also actionable for researchers. And I think
这大概就是更偏日常的那一面。另外关于研究员、关于 Anthropic 这样的研究组织,还有一层是他们对未来有一套更宏大的想象。比如说,公司刚成立那会儿,研究员就已经在讨论:怎么让 Claude 会用电脑?怎么让 AI 自己在屏幕上点来点去?所以研究员身上其实有一股我称之为「很像创始人」的劲头——特别大胆、特别有野心,而 Anthropic 这样的人一抓一大把。所以一层是对这项技术能走到哪儿的愿景;而在这个循环的另一端,是既然 Claude 已经到了用户手里,我们今天能怎么把它做得更好。所以是中长期和短期两条线并行:一大块精力用未来的视角去想,另一块盯着眼下还能改进什么。你刚才描述的其实就是我们怎么在不同版本的 Claude 上做迭代改进。我团队和研究员的合作方式是深度嵌进这些循环里,尤其是那些对用户影响很大的方向——比如视觉、computer use、写代码、agent 写代码、tool use、test-time compute 这些和用户体验直接挂钩的领域。然后去琢磨:怎么把用户反馈带回来,并且转译成研究员看得懂、也下得了手的东西。我觉得
[30:08] Dianne
that's the second piece is actually a big part of the job and sometimes a hard part of the job. So for example, we might get feedback on cla.ai. Claude hallucinated. It's very vague. If you bring that to a researcher and you say, "Please fix Claude from being hallucinated." It's not very actionable. And so part of the time of the team is understanding, okay, what's the trajectory of why that user gave that feedback? And it's like consented. And so we we we look at okay what should Claude have called tools in that moment or from its current knowledge or it called the right it looked at the right document but it looked at the wrong facts. In the first case that would have been a failure on tool use. On the second case it would have been a failure on let's say search or knowledge and search and search synthesis or it could be something around alignment. And so bring that level of detail to researchers coming up with like is this a big enough problem figure out things like evals to then describe how we've improved it like those are the levels of actionability and it's the day-to-day language of the researchers. And so we try to stay very close to how to bring that in an actionable manner uh between users to to the core model training and the research development loop.
这就是第二块,其实是这份工作里很大的一部分,有时也是最难的一部分。举个例子,我们可能在 Claude.ai 上收到一条反馈:「Claude 胡编乱造。」这话很虚。你要是直接拿去跟研究员说「请把 Claude 的幻觉问题修一下」,根本没法落地。所以团队有相当一部分时间花在搞清楚:这个用户为什么会给出这条反馈?当时完整的对话轨迹是什么?——这些数据都是用户授权的。我们会去看:Claude 当时是不是本该调用工具,还是应该直接用已有知识回答;又或者它其实调对了、文档也找对了,但引用错了里面的事实。第一种情况属于 tool use 的失败,第二种属于搜索、知识和搜索结果综合环节的失败,也可能是 alignment 相关的问题。把细节拆到这个颗粒度再交给研究员,接着判断这个问题够不够大,再设计 evals 来衡量我们到底有没有改进——这才是研究员日常语言里「可执行」的层级。所以我们努力贴得很近,把用户那端的东西以可落地的方式,送进核心模型训练和研究开发的循环里。
[31:36] Lenny
I was talking to someone the other day about how feels like research AI research is uh the place to be now if you want to be very successful in life. What does it take to become a really successful researcher from what you can tell uh you know not everyone can get in not everyone's brain is going to work this way but just say people are like hey I want to explore this career path from what you've seen what does it take to to make it there
前几天我还跟人聊,感觉现在要想这辈子混得特别成功,AI 研究就是最该去的地方。以你的观察,成为一名真正出色的研究员需要什么?当然不是谁都进得去,也不是谁的脑子都长成那个样子,但假设有人说「我想试试这条职业路径」,从你看到的情况,得具备什么才走得通?
[31:59] Dianne
researchers generally are research and product managers working with research or both
你问的是研究员本身,还是和研究团队一起工作的产品经理?还是两个都要?
[32:03] Lenny
let's do both but uh the researchers like you know PM's working researchers also going to be very successful but it feels like everyone's trying to you know poach all the top researchers across every company so just I I know you're not an AI researcher, but just from what you've seen, just like what does it take to make it in that in that career path?
两个都聊聊吧。不过重点说研究员——当然,跟研究员一起干活的 PM 也会很吃香,但现在感觉每家公司都在疯抢顶尖研究员。我知道你不是 AI 研究员,但就你的观察,在这条路上要怎么才能出头?
[32:21] Dianne
Yeah, I think a lot of the most successful researchers and research leadership at Anthropic are folks who are really strong first principles thinkers about problems. Like they reason through problems really well. um who are just passionate about their research area and have a bold description of what that could look like and then who are actually close to the details and so uh you know our like leadership our chief scientists our heads of like fine-tuning and like RL folks are actually really close to the training runs and actually look at things like how the training run is eval looking at the underlying data. So like actually staying really close and be excited to be in the details I think have been like a sign of like really strong researchers and developing taste. And I think like another piece is just like their ability to think big over time and be like very ambitious, right? like the Dario like we can transform software engineering and and and the and I think uh going in that direction you learn so much you get you had to shoot for the stars in in in many ways across u your ideas I think in order to be a a successful researcher
Anthropic 这边最成功的那批研究员和研究负责人,通常是特别强的第一性原理思考者——面对问题他们能一层层推理下去。他们对自己的研究方向是真的有热情,而且能把「这件事最终能长成什么样」讲得很有野心。同时他们又贴近细节:我们的领导层、首席科学家、负责 fine-tuning 和 RL 的几位,是真的会盯着训练跑得怎么样、evals 表现如何、底层数据长什么样。所以能扎进细节、并且享受待在细节里,一直是顶尖研究员的标志,品味(taste)也是这么练出来的。另一点是他们敢于把事情长期地想大、极有野心——就像 Dario 说的,我们可以彻底改变软件工程。我觉得朝那个方向走,你会学到特别多;很多时候你的想法就得往星星上射,才有可能成为一个成功的研究员。
[33:48] Lenny
I I love just this meme of just be more ambitious comes up so often now which is so hard like it's it's easy to say that it's hard to actually just like how big can you and how that's so much of what AI now unlocks. Just be more ambitious.
我特别喜欢现在到处都在冒的这个梗——「把野心放大一点」。说起来容易,做起来太难了,难就难在你到底能把事想多大。而这恰恰是 AI 现在解锁的东西:你可以更有野心。
[34:01] Dianne
Yeah.
对。
[34:02]
Yeah.
嗯。
[34:03] Dianne
I think it's thinking through it once or twice and to end and then being I think stubborn about the uh area and maybe more uh loose around the exact like approach. Um it it is a question we challenge ourselves with. uh but the technology is moving so quickly and so how do you make sure what you're building is actually uh forward compatible and so it's also actually part of like I think the core product development loop to think bigger right uh one thing I ask the team frequently or how I think about when we're building a product is let's say claude 8 comes around what do what changes in what users do and then what should what does that mean for how you're building today? Is it going to be forward compatible to that experience, right? So like just grounding it's I think um being ambitious is very broad and so trying to like ground it in in some ways of describing describing that
我觉得关键是把这件事从头到尾想通一两遍,然后在「方向」上很固执,在「具体怎么做」上保持松弛。这也是我们经常拷问自己的问题。技术跑得太快,你怎么确保今天在做的东西是向前兼容的?所以「想得更大」其实是核心产品开发循环的一部分。我经常问团队,或者说我做产品时会这样想:假设 Claude 8 出来了,用户的行为会发生什么变化?那这对你今天该怎么搭这个产品意味着什么?它能不能兼容到那个未来的体验?「有野心」这个词太宽泛了,所以我们试着用这种方式把它落到实处。
[35:12] Lenny
and also yeah everything heading in a direction that all is cohesive and makes sense versus just ambitious in a completely different direction. Speaking of ambition and cloud8, uh, Fable Mythos recently feels like hit this very new kind of tipping point with models where it used to be you have an awesome model, release it. Hey everyone, welcome. Opus45 is out, everyone can use it. Mythos went in a very different direction. It got blocked. There was a lot of scrutiny, a lot of concern about what it was capable of. Uh, all the companies had to go make sure it wasn't going to hack into all their systems. And it feels like now every model because they continue to get better will now have a lot more scrutiny and there will be more restrictions on who can use them which feels like a big deal. How do you think about that? How does that change the way you operate?
对,而且要保证所有东西朝一个连贯、说得通的方向走,而不是各自野心勃勃地往完全不同的方向跑。说到野心和 Claude 8——最近 Fable Mythos 感觉把模型带到了一个全新的临界点。以前的路数是:模型做好了就发布,「大家好,Opus 4.5 上线了,人人可用」。Mythos 走的完全是另一条路:它被拦住了,受到大量审视,各方都在担心它到底有什么能力,所有公司都得先确认它不会黑进自己的系统。感觉从此以后,因为模型只会越来越强,每一个模型都会被更严格地审视,能用的人也会受到更多限制——这事挺大的。你怎么看?它会怎么改变你们的做事方式?
[36:02] Dianne
I'm going to maybe leave the policy and the export control side to to folks that um own that and work on that. Um I think the product question and how we interact with these internally is I think as you mentioned as frontier models become more capable the safeguards and the ways of red teaming and testing and the pre-release process uh also needs to evolve and adapt quickly to to address that. And so one example is you know before fable models we didn't have as strong of let's say fallback UX's and systems because our our our goal was to make sure that like there is asymmetrical benefit for this technology and to minimize like the downside or like a severe risk of of it. And so we ended up building like fallback systems so that users will still get a great response from Opus 4. And so I think there's a piece around uh as we evolve and like improve safety systems. How do we continue to develop and deliver great user experiences?
政策和出口管制那一块,我还是留给专门负责的同事来讲吧。从产品角度、以及我们内部怎么应对来说:正如你提到的,前沿模型能力越强,安全防护、red teaming、测试和发布前流程也得跟着快速演进和适配。举个例子,在 Fable 系列模型之前,我们的兜底 UX 和兜底系统没有那么强,因为我们的目标是确保这项技术带来的收益是非对称的,同时把下行风险、尤其是严重风险压到最小。于是我们搭了一套兜底系统,让用户即便碰到限制,仍然能从 Opus 4 拿到很好的回答。所以这里有一条主线:在我们不断演进和加强安全系统的同时,怎么继续做出并交付优秀的用户体验。
[37:15] Dianne
I think there's more that we can do on both sides. And so you'll see us innovating, improving on what we call now the model safeguards package uh more and more in the coming coming weeks and months.
我觉得这两边我们都还有很多可以做的。接下来几周、几个月,你会看到我们在现在称为「模型安全防护包」(model safeguards package)的这套东西上持续创新和改进。
[37:29]
What's really interesting and just like unexpected here is creates this really interesting advantage for anthropic where you have access to the latest stuff and this is going to happen at every lab. Everyone's going to keep improving and it's it creates this unfair advantage within the labs to have access to the best stuff that other people can't yet outside of your control. You'd prefer everyone use it. So it's a really interesting this new feedback loop that's going to start where models that are so advanced are only accessible to certain companies and that's going to be a whole new unexpect it's like a second order effect of of all these restrictions. Our goal is to be uh to develop these systems and the models to be as inclusive as possible. Um I think our goal is to not have that happen uh for the general purpose general use like technologies and to make it more accessible. I think, you know, it this is like one of our top priorities right now to kind of reduce what we're seeing there.
(Lenny)这里有个特别有意思、也挺出乎意料的点:它反而给 Anthropic 造出了一种优势——你们自己能用上最新的东西。而且每家实验室都会这样,大家都会持续变强。于是各家实验室获得了一种并不公平的优势:手上有别人暂时用不到的最强的东西,尽管这并非你们所愿,你们其实希望人人都能用。所以这会催生一个很有意思的新反馈循环:最先进的模型只有少数公司能碰到,这算是所有这些限制带来的二阶效应,会是一个全新的局面。(Dianne)我们的目标是把这些系统和模型做得尽可能普惠。对于通用型的技术,我们并不希望出现你说的那种局面,而是希望它更可及。这其实是我们现在最优先的事项之一,就是要把已经看到的这种苗头压下去。
[38:20] Lenny
Yeah, that makes sense. I would imagine you'd want as many customers if people using this thing as possible. This episode is brought to you by Mercury, radically different banking, loved by over 300,000 entrepreneurs and now with command. I've been a customer of Mercury's for over 6 years. I have never once thought about leaving. Mercury is basically what happens when banking is built by product people, not by bankers. They make it so easy, dare I say fun, to send invoices, move money [music] around, set up virtual cards for folks on my team. Does your bank have an API, a terminal native CLI or an AI ready MCP server? I don't [music] think so. And just recently, they launched Command, a conversational interface built directly into Mercury, which acts as your financial operator. I've been using command to transfer money around to figure out what categories I've been spending the most money in, analyze my cash flows, and just today I used it to find out how much I've made from a specific sponsor over the past year. I just ask, [music] "How much have I made from X over the past year?" 10 seconds later, I have an answer. It is so freaking cool. Visit mercury.com to learn more and apply online in minutes.
是,这说得通。我想你们肯定希望用的人、客户越多越好。本期节目由 Mercury 赞助——一种彻底不同的银行体验,超过 30 万创业者在用,现在还上线了 Command。我自己做 Mercury 的客户已经六年多,从没有一刻想过换掉它。Mercury 基本上就是「银行由产品人而不是银行家来做」会长成什么样:开发票、转账、给团队成员开虚拟卡,都做得特别顺手,甚至可以说好玩。你的银行有 API 吗?有终端里的原生 CLI 吗?有给 AI 用的 MCP server 吗?我猜没有。最近他们还发布了 Command——直接内建在 Mercury 里的对话式界面,相当于你的财务操盘手。我一直用 Command 转账、看自己的钱主要花在哪些类别上、分析现金流;就在今天,我还用它查了过去一年某个赞助商一共给我打了多少钱。我只要问一句「过去一年 X 给我付了多少」,十秒后答案就出来了。真的酷得不行。想了解更多可以去 mercury.com,几分钟就能在线申请。
[39:30]
Mercury is a fintech company, not an FDIC insured bank. banking services provided through choice financial group and column NA members FDIC. I want to talk a little bit about how the product role is changing and who who is doing well in this new world uh now that AI is such a core part of uh of our life. When you're hiring PMs, product people, when you're looking at people that do well in today's world, what are some things that you notice? What are you looking for more most? What are you looking for more? What's kind like trending up in what you find is important and what's kind of trending down? We actually on my team have not changed our hiring loop uh for three years now. Um so what we actually look for and the traits and how we evaluate uh generalists like PM's generalist like research product managers have actually been the same. Um so I think some of those traits number one is first principles thinking and this is really uh rather than pattern matching what you used to do in let's say consumer product or B2B SAS um but actually figuring out in this moment for this user group with this technology what what is the user value
(Lenny)Mercury 是一家金融科技公司,不是 FDIC 承保的银行;银行服务由 Choice Financial Group 和 Column N.A. 提供,二者均为 FDIC 成员。我想聊聊产品这个角色正在怎么变化,以及在 AI 已经成为我们生活核心的今天,什么样的人做得好。你们招 PM、招产品人的时候,看今天做得好的那些人,你会注意到哪些特质?你现在最看重什么?什么在变得更重要,什么在贬值?(Dianne)其实我们团队的招聘流程三年没变过。我们看重的特质、评估通才型 PM 和研究向产品经理的方式,一直是同一套。第一条是第一性原理思考——不是把你在消费级产品或 B2B SaaS 里那套模式直接套过来,而是真的去想:在此时此刻、面对这群用户、用这项技术,用户价值到底是什么。
[40:50] Lenny
is there an example that a lot of people hear first principles thinking they're like yes I about it. I'm good at this. What is what's an example of someone having really demonstrated really good first principles thinking?
能不能举个例子?「第一性原理思考」这话大家都听过,人人都觉得「对对对,我很擅长」。什么样的表现才算真正展示了很好的第一性原理思考?
[40:59] Dianne
I think one example is I think you think of a product manager as I own product strategy and delivering user value as but I demonstrate day-to-day by writing a PRD or writing a product vision doc. And for for my team as like research product managers, the way to drive user value is to figure out the right user feedback, the evals, right? That then can be a personification of that user need. So like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? because in order to deliver that user value uh it's not that exact artifact that people used to write in the last like one to two decades it's a new way of working and so the first think principal thinking would be let me figure out what is the thing I should do to achieve my goals rather than here is a set of activities that I've done and therefore I will continue to do
举个例子:一般大家理解的产品经理是「我负责产品策略、交付用户价值」,但日常的体现方式是写 PRD、写产品愿景文档。而对我们这种研究向产品经理来说,推动用户价值的方式是找对用户反馈、做对 evals——evals 才是那份用户需求的具象化。我们当然也写一些产品文档和 PRD,但团队里有句话叫「evals 就是新的 PRD」。因为要交付那份用户价值,靠的已经不是过去一二十年大家习惯写的那种产物了,而是一种新的工作方式。所以第一性原理思考就是:我先想清楚为了达成目标到底该做什么,而不是「我以前一直干这些事,所以我接着干」。
[42:10] Lenny
so the idea here is used to be have kind of an idea create a PRD talk to people about it. Align on the plan, design it, build it, ship it, see how it goes, iterate. What I'm hearing here is it's like, okay, here's some feedback about something that's wrong or an opportunity. Step one is the eval is now how you define what the work is versus a PRD.
所以这里的意思是,以前的流程是:有个想法,写 PRD,跟大家讲,达成共识,设计,开发,上线,看效果,再迭代。我听你说的是:现在拿到一条「哪里不对」或者「这里有机会」的反馈,第一步是做 eval——用 eval 而不是 PRD 来定义这活儿到底是什么。
[42:32] Dianne
Maybe maybe step one would be uh understanding the user painoint. And so the way to even access that user painpoint is different, right? In the past, we might do a user interview and I think if you go like deep enough, you you might have the user walk you through their user flow, the pixels. Here, you have to sweat the tokens as much as you sweat the pixels. And so, one activity we have on the team is reading the transcripts and understanding uh what was the trajectories that failed very deeply to then say was this like a hallucination? was this claw being overconfident. So like the theme of the failure actually has a lot of nuance and then that allows you to build a description a like sustained description of that painoint. Uh so that could be essentially in a new eval and is the eval on distribution right is it capturing both the positive situations where this is failing and also areas when it should actually not fail and then bring that back to let's say research so then we can make the improvements and actually measure the quality of okay when we have opus 5.5 is this area improving or not is claude now able to uh identify the right places in the document uh and pull the right synthesis out. So it's just the actionability like and and shortening the distance to actionability um for for our stakeholders and partner teams like researchers um to take action on.
第一步大概还是理解用户的痛点。但现在连「怎么触达这个痛点」的方式都变了。过去我们可能做用户访谈,如果挖得够深,会让用户带你走一遍他的操作流程,抠到像素级。而现在,你得像抠像素一样去抠 token。所以我们团队有一项日常动作,就是读对话 transcript,非常深入地去看那些失败的 trajectory 到底怎么回事——这是幻觉?还是 Claude 过度自信?失败的类型其实有很多细微差别。搞清楚之后,你才能把这个痛点稳定地描述出来。而这个描述本身,基本上就能变成一个新的 eval。接着要看这个 eval 是不是 on distribution——它有没有同时覆盖到确实会失败的情况,以及那些本来就不该失败的情况。然后把它带回给研究团队,我们才能真正做改进、并且能衡量质量:比如等 Opus 5.5 出来,这一块到底有没有变好?Claude 现在能不能在文档里定位到正确的地方、抽出正确的综合结论?说到底就是可执行性——把「发现问题」到「能动手改」之间的距离缩短,让我们的 stakeholder 和合作团队、比如研究员,能直接据此行动。
[44:16] Lenny
Is there an example of something like this where you found an issue or opportunity and then wrote the eval? And what is what is the eval looking like in in most cases? uh what when people want to picture an eval what is that what is what do they picture?
有没有一个具体例子——你们发现了某个问题或机会,然后写了个 eval?大多数情况下这个 eval 长什么样?大家想象 eval 的时候,脑子里该浮现出什么画面?
[44:29] Dianne
We actually uh pioneered this concept within anthropic. So uh one of the early examples is the early cloud models were not very good at following specific schemas. So like things like outputs and JSON and uh now that is fundamental to claude being able to be a good agent. Right? if you can't output a certain format, you don't know how to like access APIs, you can't call tools, etc. And so the initial uh end to end was I was hearing feedback around you know claude 2 days claude was not very good at following instructions. So then digging in with users, what do you mean by claude is not good at following instructions? Give me what situations this was happening like what's the exact like paragraph? what did you ask? What was Claude's response? Going to like that level of detail. And what I saw was something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON.
这套做法其实是我们在 Anthropic 内部最早跑通的。早期的一个例子是:早期的 Claude 模型不太会遵循特定的 schema,比如按 JSON 格式输出。而这恰恰是今天 Claude 能当好一个 agent 的基础——你如果输出不了指定格式,就不知道怎么访问 API、没法调用工具,等等。当时的起点是我听到用户反馈说,Claude 2 那会儿 Claude 不太会遵循指令。于是我追着用户问:你说「Claude 不听指令」具体指什么?给我讲讲是在什么场景下发生的,到底是哪一段?你问了什么?Claude 又回了什么?细到这个颗粒度。结果我发现,早期大家说的这类失败里,大概 80% 其实是 Claude 没写出正确的 JSON。
[45:35] Dianne
And so then, okay, let's generate maybe to start just 30 to 40 examples of when Claude was not doing this thing correctly. And then that actually is your eval set. And you could have essentially uh a prompt and a response. And if that is not working uh in the right golden answer that you might have, then that means that the the eval essentially uh is beneficial because it's identifying a painoint consistently. And so then we added that to our um repositories for evals. And when we have uh versions of claude, we actually run that eval and just check. I think at this point it's always 100% or like 99.9. And so it's no longer a pain point. Uh but in the early days was taking the user feedback, figuring out actually what they mean, can we reproduce it, is it consistent, is it a big issue, and then figuring out how to uh standardize it in a way that can be consumable for researchers. It's basically test-driven development for PMs is is the world we're living now. Uh where you write the test first. So is this just a core part of the product management job now at Enthropic writing bells?
那接下来就是:好,我们先攒 30 到 40 个 Claude 做错这件事的例子。这批例子本身就是你的 eval 集。它基本上就是一个 prompt 加一个 response,如果对照你手里的标准答案(golden answer)跑不通,那说明这个 eval 是有价值的——它能稳定地把这个痛点暴露出来。于是我们就把它加进 eval 仓库。之后每出一个新版本的 Claude,我们都会跑一遍这个 eval 检查。到今天基本永远是 100%,或者 99.9%,已经不再是痛点了。但在早期,做法就是:接住用户反馈,搞清楚他们真正想说的是什么,能不能复现,是不是稳定出现,是不是个大问题,然后把它标准化成研究员能直接消费的形式。 Lenny:这基本上就是 PM 版的测试驱动开发嘛,我们现在活在这样一个世界里——先写测试。所以写 eval 现在算是 Anthropic 产品经理工作的核心组成部分了吗?
[46:52] Dianne
I think so. I I also think it's um something I've talked to other Piana other companies about and I think it's also more and more of the skill set more broadly because a lot of the products that we're building is at the intersection of models with harnesses with a set of contexts for a set of users. And so having things like eval isn't is a way not just for uh folks working on models but generally within product uh to to get to better user experiences because you can't improve what you can't measure and a lot of this is very still tactile based. It's still very judgment based and so you have to stay close to the details
我觉得是的。我也跟其他公司的 PM 聊过,我认为它正在更广泛地成为一项必备技能。因为我们现在做的很多产品,都处在「模型 + harness + 一组面向特定用户的上下文」的交叉点上。所以 eval 这类东西,不只是做模型的人才用得上,整个产品侧都用得上——它能帮你做出更好的用户体验,因为你没法改进你衡量不了的东西。而且这件事到今天仍然非常靠手感、非常依赖判断,所以你必须贴着细节。
[47:38] Lenny
and also very non-deterministic which is a big part of this just like it's not going to give you the same answer every time. So you got to describe it kind of more broadly. It's not going to be yeah an exact match. So this is a really interesting change in the way product happens and will happen is eval is is a big part of this. Do you guys still do PRDS? Is there still like a one pager describing a problem or is it play? Okay, now you're shaking your head. Yes,
而且它还是高度非确定性的,这也是很关键的一点——同一个问题,它未必每次都给你一样的答案。所以你的描述得更宽泛一些,不可能是精确匹配。所以这是产品工作方式上一个特别有意思的变化,eval 会成为其中很重要的一环。那你们现在还写 PRD 吗?还有那种描述问题的一页纸吗,还是说……好,你在点头——所以答案是还写。
[48:02] Dianne
we we we are we do I think um when there's a very defined problem I think things like eval might be almost a shorthand. I think there's other cases where PRDs are really valuable. Um, PRDS are great vehicles for getting a very large group of people aligned on a set of sources of truth about experience and setup goals. So when we do have a model, we actually for every model we do have a PRD less necessarily for our researchers but more for our growing product surfaces, for our engineering teams, for our um stakeholders like uh legal and safety and others as just a source of truth of putting together what we're aiming to achieve so that a big group of people can row in the same direction. The other place where I do think PRDS are valuable are on the more ambiguous problems and opportunities right so we if we haven't shipped a thing like computer use we don't necessarily have a set of like user specific pain points always and I think there's value in the product vision portions of a PRD to explore what could even if a technology is not yet ready to work for everyone how do you get it to work well for some group.
我们确实还写。我觉得在问题已经定义得很清楚的时候,eval 这类东西几乎可以当成一种速记。但也有不少场景 PRD 非常有价值。PRD 是个很好的载体,能让一大群人在「体验是什么、目标怎么定」这套事实来源上对齐。所以每做一个模型我们其实都会写 PRD——倒不一定是给研究员看,更多是给我们不断扩张的产品面、给工程团队、给法务和安全这些 stakeholder 看,作为一个事实来源,把「我们要达成什么」摆清楚,让一大群人能朝同一个方向划船。另一个我觉得 PRD 有价值的地方,是那些更模糊的问题和机会。比如 computer use 这种我们还没发布过的东西,我们手上不一定有一堆现成的具体用户痛点。这时候 PRD 里「产品愿景」那部分就有价值了——去探索:哪怕这项技术还没成熟到人人可用,你怎么让它先在某一小群人身上跑得很好。
[49:25] Dianne
So you can explore the value, you can actually bring something that is uh coherent to a user group. So we do have PRDS. Um I think the application is a little different now.
这样你才能把价值探索出来,才能给某个用户群交付一个真正连贯的东西。所以 PRD 我们是有的,只是用法跟以前不太一样了。
[49:39] Lenny
Okay, this is great. There's I just had a uh Andrew for he's the head of the codeex app at OpenAI and he's you guys are aligned. Uh PD is not dead. Still very useful for specific projects and ideas. Uh great. Okay, we've closed closed the book on purity is still kicking. Okay, so we've been talking a bit about just what kind of skills are kind of emerging for product people. Um, is there anything else that you find is shifted in what patterns uh are common across people that are doing well in this new AI world in terms of product managers and folks on the product teams? Is there anything else that you're like, okay, does something you got to shift or something you look for more people? I think maybe specifically uh for folks who might be midc career or folks who have been more in a managerial like product like leadership seat. Um, one thing that I think I feel pretty strongly about is in order to be good managers of teams and PMs working with this technology, you have to be really hands-on yourself and have spent not just time tinkering but actually shipping with this technology and and and again being in the details and sweating the tokens along with your PMS and your engineer.
太好了。我前不久刚聊过 OpenAI 那边负责 Codex 应用的 Andrew,你们说法完全一致:PRD 没死,对特定的项目和想法依然很有用。行,「PRD 还活着」这一页我们就算翻过去了。 刚才我们聊了一些产品人正在浮现出来的新技能。除此之外,你还观察到哪些变化——在这个 AI 新世界里做得好的产品经理、产品团队成员,身上有哪些共同的模式?还有什么是你觉得「这个必须变」,或者你在招人时会更看重的? Dianne:可能特别针对职业中期的人,或者一直坐在管理岗、产品 leader 位置上的人。有一点我态度挺明确的:要想带好团队、带好做这项技术的 PM,你自己必须非常 hands-on,而且不只是花时间玩一玩,是真的用这项技术发过东西——同样扎进细节里,跟你的 PM、你的工程师一起去抠 token。
[51:02] Dianne
and your teams. And so even for folks that I hire who have more tenure PM experience, the onboarding plans are exactly the same as somebody who is like more uh early career and it's around understanding users, reading like consented user feedback, talking to customers. I think there's something around uh being able to like understand what to do with this, what what good looks like and having developed that in a very hands-on manner. That's important. Um it's not necessarily easy for someone to uh agree or be able to see what a what a good or great AI product or AI feature could look like if they haven't kind of experienced building themselves. Um, so I think I think there is a I I I do feel pretty strongly that like, you know, if you're a manager, you have to be hands-on. You have to spend a portion of your time actually shipping. You you have to kind of walk in the shoes of your teams. uh and and that's I I always try to carve out a portion of time uh to to actually like own one to two work streams when we have models in order to keep like keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving uh uh so I can help the team make make decisions and and make better decisions.
所以哪怕是我招进来的、PM 经验很资深的人,他们的 onboarding 计划跟一个更偏早期职业的人完全一样:理解用户、读那些经用户授权的反馈、跟客户聊。我觉得关键在于,你得知道拿这项技术能做什么、什么才算好,而这种判断力必须是亲手做出来的。如果一个人自己没经历过从零把东西建起来,他不一定能认同、甚至不一定看得出来一个好的、乃至优秀的 AI 产品或 AI 功能该长什么样。所以我确实挺坚持这一点:如果你是管理者,你必须 hands-on,必须拿出一部分时间真正去发东西,你得走一遍团队走的路。我自己也一直刻意留出一部分时间,在做模型的时候亲自认领一到两条工作流,为的是保住我对模型的 theory of mind,保住我对「模型在往哪走、进步有多快」的手感,这样我才能帮团队做决策、做出更好的决策。
[52:35] Lenny
So, what I'm hearing here is if you're not, no matter where you are in the ladder of hierarchy at a company, if you're not building yourself, if you're not actually talking to Claude, talking to Codex, building stuff, you're not going to make it.
所以我听到的是:不管你在公司层级里处于什么位置,如果你自己不动手做东西,不真的去跟 Claude、跟 Codex 对话、把东西建出来,你是撑不下去的。
[52:46] Dianne
And you should have fun working with his technology. I think that's the other piece. I think the folks that would be most successful regardless of their level are people who love working with AI and and are exploring and experimenting and carving out the time not just for the experimentation but actually hands-on shipping end to end getting the user feedback I think has to be fundamental for everyone.
而且你还得享受跟这项技术打交道的过程,我觉得这是另一半。我认为不管什么级别,最成功的那批人,都是真心喜欢跟 AI 一起工作的人——他们在探索、在做实验,并且愿意腾出时间,不只是做实验,而是真的端到端亲手把东西发出去、拿到用户反馈。我觉得这对每个人都是根本性的。
[53:12] Lenny
I 100% know what you mean there. Just like me sitting on my newsletter and this podcast just talking about stuff and like yeah that sounds great. Like every time I actually build something and I tinker with all kinds of little projects, you're just like, "Okay, I see what's happening here." And you just get so much more, it's like hard to exactly describe what you're what you what you experience actually working with the models and building stuff, but it's like a whole different world of like, "Okay, I see. Here's where the here's what they're talking about computer use. Here's what they're talking about with this limitation, this UX situation."
我百分之百懂你的意思。就像我坐在这儿写 newsletter、录播客,光是聊这些事,会觉得「嗯,听起来挺好」。可每次我自己真的动手做点什么、去捣鼓那些乱七八糟的小项目,感受完全不一样:「哦,我看明白这里在发生什么了。」你得到的东西多太多了。那种真正上手用模型、做东西的体验很难精确描述,但它就像另一个世界:「哦,原来他们说的 computer use 是这个意思,他们讲的那个限制、那个 UX 问题是这么回事。」
[53:39] Dianne
Yeah.
嗯。
[53:39] Lenny
So, yeah. So, it's just like, and you made this really interesting point that you have to have fun with it, which is not easy for a lot of people because they're pushed to use AI or they just don't know exactly what to do with it. For people that are just like, I don't know, it's just so annoying. I just have to do this. I don't know what's so like, I hate this freaking thing. Why do I have to work with this? Things are changing so much. I'm tired. Uh, advice for helping people find that find that joy in this work. I think maybe I'll reemphasize something I said earlier around just that experimentation is not an individual sport. Like some of the moments where I think I've touched practically every version of research models across 20 versions of production clause at this point and I think part of the joy comes from seeing other people discover use cases too. And so maybe one idea here would be pairing with somebody who is excited and seeing what on a use case that you care about and and working together versus um uh identifying or trying to figure out the perfect use case yourself because that might feel like work. Working with others feels like joy a lot of the time.
所以是这样。你刚才提到一个特别有意思的点:你得从中找到乐趣。但这对很多人并不容易,因为他们是被推着去用 AI 的,或者根本不知道该拿它干嘛。有些人就是觉得:「烦死了,我不得不用这玩意儿,我讨厌这破东西,凭什么我非得跟它打交道,变化这么快,我真的累了。」对这些人,有什么建议能帮他们找到那份乐趣? Dianne:我可能再强调一遍前面说过的那句:做实验不是一项个人运动。我到今天差不多摸过每一个版本的研究模型、二十来个版本的正式 Claude,但我觉得乐趣有一部分是来自看别人也发现了新用法。所以一个具体建议是:找一个已经很兴奋的人结对,围绕一个你真正在意的用例一起搞,而不是自己一个人去找那个「完美用例」——后者感觉像工作,而跟别人一起搞,很多时候感觉像快乐。
[54:56] Dianne
And is there more that we could do to bring that bring other people along? That's something like a lot of times internally we have somebody who is like very curious and them sharing an idea of a new prototype actually brings a ton more people who are like oh I didn't know this could work now with claude and so there's just some virtuous cycles here um and and ways of yeah bring continue to have joy with with this technology.
另外就是,我们还能做点什么把更多人带进来?我们内部经常出现这种情况:某个特别好奇的人分享了一个新原型的想法,一下子就带动一大票人——「哦,原来现在用 Claude 还能这么玩」。所以这里是有正向循环的,也是持续从这项技术里获得乐趣的一种方式。
[55:22] Lenny
That's such a good point. I think that's also why Twitter's so useful for a lot of this is you see other people sharing what they've done
这个点特别好。我觉得这也是为什么 Twitter 在这件事上这么有用——你能看到别人在分享他们做出来的东西。
[55:30] Lenny
and it inspires you to come up with your own little ideas and also it's just like fun to share your own thing that you've done.
这会激发你冒出自己的一些小想法;而且把自己做的东西分享出去,本身就挺好玩的。
[55:36] Lenny
So that's a really good point just like find other people to kind of play around with and look for use cases. The thing I've also heard a lot is just find like a problem you want to solve in your life or work and just open up cloud cloud code tell it here's what I want to do and it's incredible how far you can get just with like a vague idea of a problem you want to solve. Yeah. I think it gets hard in that there's so many different things that you could try.
所以这确实是个很好的建议:找几个人一起玩,一起找用例。我还常听到的另一个说法是:从你生活或工作里真正想解决的问题出发,直接打开 Claude、打开 Claude Code,告诉它你想干什么——你只带着一个模模糊糊的问题设想,最后能走多远,简直不可思议。 Dianne:嗯。我觉得难就难在,能试的方向实在太多了。
[55:57]
Yeah.
对。
[55:57] Dianne
And so you just like narrowing in on either pairing with someone, working with somebody who who is who have a lot of joy about this technology or figuring out something that you could immediately find value. Like either of them those things allow you to go deeper rather than like more high level about too many things. Um I I find it hard to keep pace with the number of prototypes or products that are out there and so my lens has been how do I go deep in one to two of them
所以你要做的就是收窄范围:要么找个人结对,找一个对这项技术特别有热情的人一起干;要么找一件能立刻见到价值的事。这两条路都能让你往深里走,而不是浮在表面、什么都碰一点。说实话,外面新出的原型和产品实在太多了,我根本追不过来,所以我给自己的取舍是:挑一两个,自己往深里钻。
[56:30] Lenny
myself. That's uh that's so interesting you say that because that's exactly it. We just had this survey uh that I I ran with uh my colleague Noam uh asking my readers just how they're feeling about all the things going on in the tech right now and AI and uh one of the most interesting takeaways we had was uh to find that happiness is exactly what you said is go deep in a couple things versus trying to just ton of little things. find a couple things to really solve well and then go deep and that is a source because a lot of the happiness people feel is when they finally unlocked a way for AI to actually make their lives better versus just like a couple messed up broken half working things.
你这么说太有意思了,因为这跟我们发现的完全一致。我和同事 Noam 刚做过一次读者调研,问大家怎么看当下科技圈和 AI 的这些变化,最有意思的结论之一,就是你刚说的那件事——幸福感来自于在一两件事上深挖,而不是同时铺开一堆小事。找到几件事,真正把它做透、做深,这才是快乐的来源。因为人们真正感到高兴的时刻,是终于摸索出了一条让 AI 切实改善生活的路子,而不是手里攥着一堆半残的、勉强能跑的破玩意儿。
[57:08] Dianne
Yeah, it's it's um how do you go from this being a check the box, right? And so like us as product people, it's then a exercise of product prioritization of your time and your energy. And and if the goal is to experiment with joy, then how do you what are the inputs that you need for that? Um, but yeah, I I I think a lot of the um I think the secret sauce of anthropic is the culture and the bottoms of nature of how people work and this like experimenting in public. Um, and by doing that, it's very much about how to bring other people along. Um, that ends up being, I think, really valuable. Yeah, I've heard this so many times from all the labs just like no no no one's exactly sure how some of this is going to be used and a lot of it is just putting stuff out early, seeing how people use it, seeing what it's what's possible and then using that information to build the actual product to lead in.
对,关键是怎么让它不再是一件「打个勾交差」的事。我们做产品的人,这其实就是一次对自己时间和精力的产品优先级排序。如果目标是「带着乐趣去实验」,那你需要哪些输入条件?我觉得 Anthropic 的秘诀在于文化,在于大家做事那种自下而上的劲头,还有这种「公开做实验」的氛围。这么做的核心其实是怎么把其他人也带上车,最后价值特别大。
(Lenny)这话我从各家实验室听过太多遍了——没人能笃定这些东西最终会怎么被用起来,很多时候就是早早把东西放出去,看人们怎么用、看能长出什么,再拿这些信息去打造真正的产品。
[58:10] Dianne
Yeah. Yeah.
对,没错。
[58:11] Lenny
I'm curious how kind of on this thread of finding ways AI for AI to help you in your work in life. Are there any interesting ways you've been using Claude lately in your work as a as a PM? I think there's a lot of things with um you know fable and things like tag. So there there I think tag is um in in the very like early days I think there's something around how you work in a different paradigm of allowing this an agent to go off and work and then bring back uh product and experiences to you. I think one area that it's not very recent, but one that um I bring up a lot with the team and I think we could do more on using AI is just like how to use it to also be more uh to have better conversations with each other to be better managers. I don't think it's necessarily uh just about raising the IQ of like experiences we build, but also I used it a lot and actually like prepping for how to have better conversations um in the moment during like crucial conversations. So, I love that book and so I actually have a skill that helps me figure out am I having am I going in the right level of detail given the the situation at hand and actually helping me be a better manager and better supporter for the team. Um, so for for like managers on the team, that's actually a thing that I've been sharing more with with a uh with our managers of okay, how how do you actually use use claude to to to make you a better coach
顺着「找到 AI 帮上工作和生活的方式」这条线,我挺好奇:你最近在 PM 工作里有没有什么有意思的 Claude 用法?
(Dianne)有不少,比如 Fable,还有 Claude Tag 这类东西。Tag 还处在很早期,但我觉得它代表了一种新的工作范式:让 agent 自己跑出去干活,再把成果和体验带回来给你。另外还有一块,不算很新,但我在团队里经常提、而且觉得 AI 还能做得更多的,就是拿它来把彼此之间的对话谈得更好、把管理者当得更好。我不觉得 AI 只是用来提升我们做出来的产品体验的 IQ;我自己就大量用它来准备那些关键对话该怎么谈。我特别喜欢《关键对话》这本书,所以我专门做了一个 skill,帮我判断:面对眼下这个情境,我切入的颗粒度对不对?它实实在在让我成为一个更好的管理者、更好的团队后盾。所以对团队里的管理者,我现在会更多地分享这个用法:怎么真正用 Claude 让自己成为一个更好的教练。
[59:50] Dianne
because it's hard sometimes to find the right perfect words and the models have a lot of perfect and right words and
因为有时候你很难找到那个刚刚好的措辞,而模型手里有大把恰到好处、分寸得当的表达。
[59:57] Dianne
uh I think there is something about how how it can actually augment us from like an ET perspective in addition to you. Oh man, there's so much interesting stuff there. So just to understand what you're doing there. So you built a skill. You're just like Claude build a skill pulling in lessons from Crucial Conversations the book which it knows enough about. You don't have to even give it the content. And then you use that skill to talk to Claude. Hey, I have this very difficult conversation coming up with a colleague.
我觉得除了智力层面,它其实也能从情商(EQ)的角度增强我们。
(Lenny)天哪,这里面有太多有意思的东西了。我先确认一下你具体是怎么做的:你做了一个 skill,就是让 Claude 把《关键对话》这本书里的方法提炼进去——它对这本书本来就足够熟,你甚至不用把内容喂给它。然后你就用这个 skill 跟 Claude 聊:嘿,我马上要跟某位同事进行一场很难的对话。
[1:00:23] Lenny
Give me some tips on how to approach it.
给我一些建议,我该怎么切入。
[1:00:25] Dianne
Yeah. And it's it's a great uh it's almost like uh coaching like individualized personalized coaching of just how to make you and and there's so much context switching that we do all day and having like Claude help me pair and help me and maybe there are times where I end up not using suggestions from Claude. Uh but it actually is uh ends up being very helpful for for just coming up and brainstorming. Am I thinking about reactions in the right way? How do I actually uh go a bit deeper faster? Build trust faster, uh be more direct.
对。这体验非常好,几乎就像有一位私人化、个性化的教练。我们一整天要在无数上下文之间来回切换,有 Claude 陪着我一起想,真的很有帮助。也有些时候我最后并没有采纳 Claude 的建议,但光是拿它来做头脑风暴、把思路先跑一遍,就已经很值了:我对对方反应的预判合理吗?怎么能更快地聊到深处?怎么更快建立信任?怎么把话说得更直接一些?
[1:01:05] Lenny
Yeah, man. I have so many questions here. This so interesting. Uh one is just like there's concern people are going to start talking the way AI writes because they're talking AI so much and it's going to be like Diane, it's not this, but it's that. Uh I know that you're not doing that, but that's a you know, a concern people have. Let me just ask about that, I guess. Do you fear this? There's this, you know, brain rot atrophy stuff people talk about it. We're just so reliant on AI now and we stop learning and thinking and, you know, overly AI thoughts on that being so close to it and being so integrated with with AI constantly.
天,我这儿问题太多了,实在太有意思。一个是:有人担心,人跟 AI 聊得太多,说话方式会开始变得像 AI 写出来的——比如「Dianne,这不是……,而是……」那种腔调。我知道你没有这个问题,但确实有人在担心。那我就直接问吧:你怕吗?外面常说的「脑子退化」「思维萎缩」,说我们太依赖 AI,慢慢就不学习、不思考了。你天天泡在 AI 里、跟它高度绑定,怎么看这件事?
[1:01:38] Dianne
A lot of actually thinking process and writing process are tied together for me personally. And so I think there are ways where I use claw to augment my thinking. But what I want to make sure and maybe this is what you're describing is Claude doesn't take over all of my thinking for me. And so I think depending on the situation, depending on how much more personal judgment I want to have in a situation, I might uh um come up with my own POV first and then work with Claude through that. Um and making sure that like I maintain my sense and tone throughout. I think there are then other things like updates right we have like monthly business reviews and then in those cases it's much more I want actually want it to be standard and I want it to be much more like it gets a cris crystallized information in the right way and maybe and I have a skill and like we're augmenting and improving our skill for that but I want to get a to a place where like the monthly business review the writing of that is potentially asymmetrically less valuable than the thinking and so how do I get that piece delegated to claude fully and I'm more of a reviewer and a verifier of that information. So I think it depends on like what you're using Claude for and what you're trying to convey and like is there is there asymmetrical value in in delegating more to Claude.
对我个人来说,思考的过程和写作的过程是绑在一起的。所以我确实会用 Claude 来增强我的思考,但我要守住的一条——可能这也正是你说的那个隐忧——是别让 Claude 把我全部的思考都接管过去。所以要看场景:如果这件事我希望更多体现自己的判断,我会先自己形成一个观点,再拿着它去和 Claude 打磨,并且确保从头到尾保住我自己的语感和语气。
但也有另一类事情,比如各种进展汇报——我们每个月有 monthly business review——这种场合我反而希望它是标准化的,希望信息能被以正确的方式结晶出来。这块我也有一个 skill,而且还在不断增强和优化它。我想达到的状态是:月度业务回顾里,「写」这个动作的价值,相比「想」这个动作是不对称地低的,那我怎么把「写」这块完全交给 Claude,让自己更多地扮演审阅者和校验者的角色。所以这取决于你拿 Claude 来干什么、你想传达什么,以及——把更多东西交给 Claude,在这件事上是不是存在这种价值的不对称。
[1:03:11] Lenny
What I'm also hearing the first tip is really great which was think first have a point of view and then kind of use Claude as a sparring partner almost to evolve the idea push back on the idea.
我还听到你给的第一条建议特别棒:先自己想,先有观点,然后把 Claude 当成陪练,用它来推进想法、反驳想法。
[1:03:22] Dianne
Yeah. Yeah. And I think this is where things like actually our alignment research and safety research is helpful because it what you don't want is like a AI that just agrees with you, right? What you want is this technology to actually augment and grow and like get to a better outcome. And so sometimes it's having Claude push back makes me better. And so that's great. like a co-orker, I want somebody to push back when my ideas are not fully formed.
对。我觉得这正是我们的对齐研究和安全研究派上用场的地方——你最不想要的,就是一个只会顺着你说的 AI。你想要的是这项技术真能增强你、推着你成长,最后得到一个更好的结果。所以有时候让 Claude 来反驳我,反而让我变得更好,这挺棒的。就像同事一样,当我的想法还没成型时,我希望有人能顶我一下。
[1:03:53] Lenny
I want to hear more about that. But I've heard that when Ben man was on the podcast, he talked about the constitution that is built into Claude and how unintuitively the work and the focus on safety and alignment as you said and this constitution that describes how Claude should think and operate that actually you would think that would limit the abilities of Claude and make it less fun and interesting. It's exactly the opposite. Claude is the most interesting personality. I hear that constantly. It's just like I much prefer talking to a like open claw famously was built on claude and then people were forced to we won't get into it were forced to switch to and they're like this is so bad this is not who I'm used to talking to. Uh so that is I think a really interesting point I just want to make sure we spend a little time on. Why is it why is that the case just this focus on alignment safety having this clear constitution? Why does that make Claude better and and more interesting to talk to? Also,
这个我想多聊聊。Ben Mann 上我这档播客的时候讲过 Claude 内置的那份 constitution(宪法),以及一件很反直觉的事:正是这种对安全和对齐的投入——就像你说的——以及这份规定 Claude 该如何思考、如何行事的宪法,你会以为它是在给 Claude 套上枷锁,让它变得没那么好玩、没那么有意思,结果恰恰相反,Claude 是最有意思的那个「人格」。这话我不停地听到有人说。就像大家熟知的,OpenClaw(原名)当初是基于 Claude 做的,后来被迫——具体细节就不展开了——被迫换成别的模型,用户的反应是:这也太糟了,这根本不是我习惯对话的那个。所以这一点我觉得特别有意思,想专门花点时间聊:为什么会这样?为什么对齐和安全的投入、有一份清晰的宪法,反而让 Claude 更好、更有意思?
[1:04:49] Dianne
in order to make Claude as like intelligent and as capable as possible, being able to have Claude actually push back in the right points and then add it's like a yes or no and actually helps you come to a better conclusion. So, I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude? It's a little bit meta, but using a research version of Opus, asking it to figure out how it should price and being able to come out with better outcomes is a goal at the end of the day. And so having AI not just be an assistant, not just be a doer, not and being delegated task, but figuring out is it doing the right thing. That's actually very integrated with knowing when to push back,
要让 Claude 尽可能聪明、尽可能能干,它得能在该反驳的地方真的反驳你,而不只是给个「是」或「否」,这样才能帮你得出更好的结论。举个例子,我用 Claude 帮我判断过:我们对下一版 Claude 的定价决策做对了吗?这有点套娃——我用的是 Opus 的一个研究版本,让它自己想该怎么定价,最终能推出更好的结论,这才是最终目的。所以 AI 不该只是个助手、只是个执行者、只是个被派活的对象,它还得能判断「现在做的这件事对不对」。而这一点,和「知道什么时候该反驳你」是深度绑定的。
[1:05:43] Dianne
right? That's part of knowing when you should be proactive. Proactivity is not a necessarily always doing a thing that you are scheduled to do. It is knowing when to come up with a new idea. And so in order for Claw to be more useful, the general approach has to be that it knows when to push back. It's a core part of the characteristics together uh of the models.
对吧?这也是「知道什么时候该主动」的一部分。所谓主动性,并不一定是把安排给你的事做完,而是知道什么时候该提出一个新想法。所以要让 Claude 更有用,总体思路必须是:它知道什么时候该反驳你。这是模型整体性格特质里很核心的一块。
[1:06:09] Lenny
That is so interesting. It's so interesting that that is what a big part of like it be it being less compliant is almost what makes it better and more useful because we need that. Like I've had so many people where they're like, "Hey, like AI told me I was right." and like no I wish I wish to other people.
太有意思了。有意思就在于:让它「没那么顺从」,恰恰是让它更好、更有用的关键,因为我们真的需要这个。我见过太多人跟我说:「嘿,AI 说我是对的。」我心里想:不是,我真希望它当时能怼你一下。
[1:06:26] Dianne
Yeah. And it comes back to our earlier point around thinking, right? How do you protect your thinking?
对。这又回到我们前面说的思考那点了——你要怎么守住自己的思考?
[1:06:31]
Um if you have a AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you and you should come away at the end of the day having better ideas because you worked with Claude. That should be the hero goal, not just making your ideas 10% better. Yeah, I love this since like it used to be think 10x. I used to be the the way you know founders push people like what if we 10x this and I love what I keep hearing is like it's like how do we go thousandx from this idea? What is the most ambitious version of this? I want to come back to something that I I was thinking about as we were talking about uh talking to Claude constantly. Um it's very clear when AI has written something still. It's funny that it's a large language model. you would think of all things it would be very good at writing and interestingly just no AI is very good at writing it's always very clear this was AI written do you think we'll get to a place where we will not know this was AI
(Dianne)如果你有一个能当思考搭档的 AI——真正的思考搭档不会只是顺着你说,它得给你添点东西。一天下来,因为和 Claude 一起工作过,你的想法应该变得更好了。这才是该追的核心目标,而不是把你的想法改进 10%。
(Lenny)对,我太喜欢这个说法了。以前大家讲的是「想大 10 倍」,创始人推团队的方式就是「要是我们把这个做到 10 倍会怎样」。而我最近一直听到的说法是——怎么让这个想法变成一千倍?它最有野心的版本长什么样?我想回到刚才聊「随时随地跟 Claude 对话」时冒出来的一个念头:到今天,一段文字是不是 AI 写的,还是一眼就能看出来。挺讽刺的,它明明叫大语言模型,你会以为写作恰恰是它最擅长的,可结果是没有哪个 AI 写得真好,一读就知道这是 AI 写的。你觉得我们会走到那一天吗——看不出这是 AI 写的?
[1:07:32] Dianne
I think it depends on what's the uh goal that you're looking to achieve by knowing yeah uh what's the eval um I actually do think there's more that we could be doing on making Claude write better. There's actually very active efforts um on on my team and on the research side about making Claude write better. Just generally I think it should be clear where an idea is ident is being led by you or by you Lenny or me Diane. I think it really depends on uh what's the goal of that writing. like for something like a monthly business review, I would actually love to have that end to end be written by Claude. Uh,
我觉得这取决于,你想通过「知道它是不是 AI 写的」达成什么目的——也就是评判标准是什么。我确实认为,在让 Claude 写得更好这件事上我们还有很多可以做。我团队这边和研究那边现在都有非常活跃的投入,就是要把 Claude 的写作能力提上来。总体上我觉得,一个想法到底由谁主导,这件事应该是清楚的——是你 Lenny,还是我 Dianne。但具体怎么看,真的取决于那段文字的目的。比如月度业务回顾这种东西,我其实巴不得从头到尾都由 Claude 来写。
[1:08:21] Lenny
and obviously and not make it feel like it was written by a human. It's such an interesting point you're making like is it actually better for us to know that it's AI versus not.
而且显然,也不需要把它写得像人写的。你这个点特别有意思——到底是让我们知道「这是 AI 写的」更好,还是不知道更好。
[1:08:29] Dianne
Yeah. But but it's it's um but it's also for maybe the lens is more around like verifiability or who's verifying
对。不过……也许更合适的视角是「可验证性」,或者说,是谁在做验证。
[1:08:40] Dianne
the output. Right. Right. like who's signing off. Uh maybe less around who's writing, but who's verifying who's signing off. That becomes like more what matters than who's writing it.
……谁来验证这个产出。对,对。就是谁来签字背书。重点可能不在于谁写的,而在于谁来验证、谁来签字,这件事变得比「谁写的」更要紧。
[1:08:52] Lenny
Why Why do you think AI is not great at writing? Like my guess is it has studied all of the best writing in all of humanity. It's figured out here's the best way to write. And now that we and it's just there's only so many ways to to write. And so we've just recognized, okay, this is what AI does. It has these tropes. Is that the core of it? Is there something else that's keeping it from being a great writer? Ironically, being a large language model of all things, you think it'd be really great at language.
你觉得 AI 为什么写不好?我的猜测是,它把人类历史上最好的写作都学了一遍,摸清了「最好的写法」长什么样。可写法就那么几种,于是我们慢慢就认出来了——哦,这就是 AI 的路数,它有那么几个套路。核心原因是这个吗?还是说另有什么东西卡着它,让它成不了一个好作者?挺讽刺的,它偏偏是个大语言模型,你会觉得它在语言上理应特别强。
[1:09:21] Dianne
I think part of it is also uh we need to invest more in training improvements to make AI continuously strong on areas like writing. Um I think it's also like the technology is jagged edged like like we mentioned. So sometimes when the models were good at writing but not agentic our our thesis is how do we make the models more agentic or call the right tools. Now that that's improved a bit then it's well now these other areas actually become more of the rough edges. And so I think we're in one of those moments with writing where uh we need to actually just focus and prioritize on training the models to be like great at this area and like that is an active a very active area for us that you mentioned.
我觉得一部分原因是,我们得在训练上投入更多,才能让 AI 在写作这类能力上持续变强。另外就像我们前面说的,这项技术的能力是参差不齐的(jagged edge)。所以曾经有一阵,模型写作还行但不够 agentic,那我们的判断就是先解决「怎么让模型更 agentic、更会调对工具」。现在这块改善了一些,其他方面反倒就成了新的毛边。我觉得写作现在正处在这样一个节点——我们得真正把注意力和优先级放上去,专门训练模型在这块变强。这正是你刚提到的、我们眼下投入非常大的一个方向。
[1:10:11] Lenny
Okay. I'm glad I'm glad. And also uh it's going to be interesting once AI is so good we're like I don't know who wrote that but um to your point sometimes we actually want to know that it's AI. That's really interesting. I never thought of it that way. The other interesting part of this is that there's that comedian who was joking that we're like on a plane and the Wi-Fi is down and we're just like, "What the hell? The Wi-Fi is not working on this plane. The sucks. How dare you?" When you're like in a in a tube in the sky flying like a bird and uh how dare you complain that the Wi-Fi doesn't work. Like your point is there's so much advancement and so much power. Uh we can't fix it all. We can't make it all work the best possible. And so uh basically AI writing has been not the priority and it feels like there's more investment happening there.
好,那我就放心了。另外,等 AI 好到我们看着一段话说「我也分不清这谁写的」,那会挺有意思的。但按你的说法,有时候我们反而是想知道这是 AI 写的。这点很有意思,我以前从没这么想过。还有个有意思的角度:有个喜剧演员讲过一个段子——我们坐在飞机上,Wi-Fi 断了,就开始抱怨「搞什么,这飞机上 Wi-Fi 又不好使,太烂了,你们怎么敢这样」。可你人正待在一个铁管子里、像鸟一样飞在天上,居然还好意思抱怨 Wi-Fi 不通。你说的其实是同一件事——技术已经进步了这么多、能力这么强,我们不可能样样都修好、不可能每一块都做到最优。所以基本上,AI 写作一直不是最高优先级,而现在感觉这块的投入在变多。
[1:10:53] Dianne
Yeah, I think like tone and character is a priority. I think it's this advancement of the technology is a work in progress and so we made we we see a leap or emergence of like a jump in agentic behaviors and so that is a new normal and then these other capabilities needs to continue like improving
对,我觉得语气和人格(tone and character)是有优先级的。技术的推进本来就是个进行中的过程:我们在 agentic 行为上看到了一次跃迁、一次能力涌现,那就成了新的基准线,然后其他这些能力得接着往上追。
[1:11:16] Dianne
and I think once we improve let's say writing and like tone and character uh we probably will say like
我猜等我们把写作、把语气和人格这些提上去之后,我们大概又会开始问——
[1:11:24] Dianne
how do we have Claude be even more proactive like productivity is an opportunity and that's human nature like we want to make ourselves better. We want to make this technology better. Um so yeah I I think we're applying it to to AI which is the right thing. We should be making it better.
怎么让 Claude 更主动一点?生产力这块也还有机会。这其实就是人性:我们总想让自己变得更好,也想让这项技术变得更好。所以是的,我们把这股劲用在 AI 身上,我觉得是对的,本来就该一直把它做得更好。
[1:11:41] Lenny
I want to ask you a couple questions I'd like to ask folks working at the very center of the future of that is coming. Um one is where do you think human brains will continue to be most valuable over the years? I know anthropic's mission and and vision is we'll reach a GI a super intelligence. So in the future maybe nowhere but before we get there where do you think human brains will continue to be most valuable as we've approached that that timeline?
我想问你几个我常问那些身处「未来最中心」的人的问题。第一个是:接下来这些年,你觉得人脑在哪些地方会继续最有价值?我知道 Anthropic 的使命和愿景是我们终将走到 AGI、走到超级智能,所以到了那一天可能哪儿都轮不到人脑了;但在抵达之前,随着我们一步步逼近那个时间点,你觉得人脑最有价值的地方在哪?
[1:12:09] Dianne
We started to talk about making claude and models better at judgment um especially in the last um year or so. I think judgment is one and is an area where it's an accumulation of so much nuance and so much experience and these systems haven't experienced as much as humans have and so I think that hard-earned like judgment is a a a area for for product leaders and just generally um will continue to be really critical. There are so many things AIs can build. which one are the things that you know an or like lab should build right a lot of that requires like human judgment persistence so proactivity these are all traits that are beyond just general capabilities but just behaviors and characteristics of like people at that level of like how do you get to the best solutions how do you create the the best experiences so I think those types of traits are actually the tactile uh traits that I think will uh continue to be important. Um I think there is also uh still a lot of like capabilities and subject matter expertise as well. I think you know software engineering has been really transformed by AI. I think there's areas like uh biology, life sciences. These are all things that um we're just kind of at like the foot of the exponential on like maybe software engineering. We're on the exponential on some of these area other areas. We're not quite there yet. And so um I think you're seeing us ship things like cloud science investing in these areas because those are areas that um I think is just bring the this technology to society and having a positive benefit for society.
大概最近一年左右,我们开始讨论怎么让 Claude、让模型在「判断力」上更强。判断力算一个。这块是大量微妙细节和经验的累积,而这些系统经历过的事远不如人类多,所以那种一点一点挣来的判断力,对产品负责人、乃至对所有人来说,都会继续极其关键。AI 能造的东西太多了,那到底哪些是一家公司、或者一个实验室该造的?这里面很多都需要人的判断力、韧性和主动性。这些特质超出了通用能力的范畴,更像是到了那个层级的人身上的行为方式和性格——你怎么走到最优解,你怎么做出最好的体验。所以我觉得,这类特质才是真正拿得出手、会一直重要的东西。另外我觉得,专业能力和领域专长也还有很大空间。软件工程确实已经被 AI 彻底改造了,但还有像生物、生命科学这些领域——软件工程也许已经踩在指数曲线上了,而这些领域我们还只是站在指数曲线的脚下,远没到那一步。所以你会看到我们推出 Claude for Science 这类东西、往这些方向投入,因为这些才是把技术真正带进社会、给社会带来正向价值的地方。
[1:14:09] Dianne
So I think there's a lot more to go there.
所以我觉得这块还有很长的路可以走。
[1:14:11] Lenny
Another question I want to ask is um as someone with kids, how do you think about what you are encouraging them to learn? or do you think you're gonna nudge them to be successful in this wild new world that we're entering?
还想问你一个问题:作为一个有孩子的人,你怎么想该鼓励他们学些什么?或者说,在我们正踏进的这个疯狂新世界里,你会往哪个方向推他们,让他们能过得好?
[1:14:27] Dianne
I actually think it's a lot of the same traits like you and I probably grew up with, which is
我其实觉得,还是很多和你我从小被培养的一样的特质,比如——
[1:14:33] Dianne
curiosity for learning, persistence, believing in your own inner voice, developing, and then believing in your own inner voice. Like I have a four-year-old, I have a 8-year-old. It's on us to help uh it's on me to help them develop their indoor voice and whether that's being opinionated and taking a stance to me right and developing that encouraging that uh I think that those types of skill sets are things that um is important in the future and like having their own individual voice.
对学习的好奇心、韧性,还有相信自己心里的那个声音——先把它养出来,然后相信它。我家一个四岁、一个八岁。责任在我们身上、在我身上,去帮他们把内心那个声音养起来。哪怕它表现出来就是有主见、敢亮明立场——比如敢当着我的面站队。把这个养出来、去鼓励它,我觉得这类能力在未来很重要,就是要有属于自己的声音。
[1:15:10] Lenny
That is so interesting. It's so related to the answer you had when I asked about how to avoid a brain rot essentially and overrelying on AI which is just keep focused on your own point of view and your own perspective before you overly AI and just this idea you're describing of building that in kids is is really important. Uh that is so interesting and I love how this all this kind of connects judgment persistence in a point of view of your own.
太有意思了。这跟我之前问你「怎么避免脑子退化、避免过度依赖 AI」时你的回答特别呼应——就是在你转身去问 AI 之前,先守住自己的观点和视角。而你现在讲的,是要把这个东西从小在孩子身上养出来,这真的很重要。太有意思了,我很喜欢这些最后是怎么串起来的:判断力、韧性,加上属于你自己的观点。
[1:15:34] Dianne
Yeah.
对。
[1:15:35] Dianne
Both for kids and also adults.
不管对小孩还是对大人都一样。
[1:15:36] Lenny
Yeah. Anything we think about um for your
对。那轮到你自己的孩子,你会怎么考虑这件事?
[1:15:39] Dianne
Oh man. Well, like the question I'm thinking about is just when to get them on like some AI thing, you know, when I have a three-year-old, so it's pretty early for that, but you know, how do you get how do you onboard them to this crazy thing? I had I was at an event recently and bunch of parents were talking about how they think about AI in their kids and one person had a really interesting approach which is uh keep them on the very early models so that they still have to struggle a bit and not get all the answers immediately. thought that was interesting. Like an open source local model,
天哪。我现在在琢磨的问题是,到底什么时候该让孩子开始接触 AI 这类东西。我家孩子才三岁,现在谈这个还挺早的,但你到底该怎么把他们领进这么疯狂的一个世界?我前阵子参加一个活动,一群家长在聊他们怎么看 AI 和孩子的关系,其中一个人的做法特别有意思——让孩子只用非常早期的模型,这样他们还是得自己费点劲,而不是一问就立刻拿到全部答案。我觉得这个思路挺有意思的。就像用开源的本地模型,
[1:16:08]
not stable.
不太稳定的那种。
[1:16:10]
Yeah.
对。
[1:16:10]
Yeah.
嗯。
[1:16:11] Lenny
Yeah. And curiosity is something uh I I keep mentioning Ben man, but his answer actually to this question has always stuck with me, which is um curiosity and also just like he's a big fan of Monosuri, which is what I'm we're encouraging for our kids. So, there's something there. Maybe a last question just along kind of along these lines, something Fiona Fun actually suggested to ask you uh who's recently on the podcast. How do you stay just recharged and not burn out being in the center of this crazy storm of AI as a mom uh working in, you know, we're seeing the research work at Enthropic. Uh I just like we're living through the most unprecedented time working at just like being, you know, being on the outside of Anthropic. It's crazy. I don't even know what it's like to be on the inside. Um what have you learned about avoiding burnout, staying recharged, staying sane during the middle of all this? In 2024, we shipped four models for the in the whole year or four series of models and I think we did more than that volume in just Q2 of this year.
对。还有好奇心。我老是提到 Ben Mann,但他对这个问题的回答一直让我印象很深——就是好奇心;另外他也特别推崇 Montessori(蒙台梭利),这也是我们家现在在给孩子做的方向。所以这里面确实有点东西。那也许问最后一个问题,大致还是沿着这条线来的,这个问题是 Fiona 建议我问你的,她最近刚上过这档播客。你身处 AI 这场疯狂风暴的正中心,同时还是一位妈妈,还在 Anthropic 做研究这条线的工作——你是怎么让自己不断充电、不被 burn out 的?我们正在经历一个前所未有的时代,我只是站在 Anthropic 外面看都觉得疯狂,根本想象不出在里面是什么滋味。关于避免 burnout、保持精力、保持清醒,你都学到了些什么?2024 年我们一整年发了四个模型,或者说四个系列的模型,而我感觉今年光 Q2 发的量就超过了那个数。
[1:17:13] Dianne
[laughter] I think I've been really lucky with uh the team that we grown and built both the stakeholders on the research side and within our research product management team. Um I think that one of the magical parts about approaching all of this is that it's not an individual sport. Um there's like a sense of radical ownership and team collaboration that I think sometimes it does feel like a high performance sport because you're in very critical decisions. there's new information about users about training and you have to make recommendations and judgments and decisions very quickly and nobody can do that sustainably by themselves. Um, and so I think what's really helped is having a team that is incredible, who looks out for each other, who, you know, night before a launch, even if they're not the core DRRi on that model, will stay up and help the DRRi, who uh to review the blog post and make edits and come up with better demos and knowing to be each other's sort of extra hand. I think it's very easy if you take all of this change on your own shoulders to feel like you're alone and to feel like you have to do everything. Uh but I think one of the like magical parts of anthropic is this ability for us to uh figure out what are those opportunities to help each other and actually then taking the next mile of like mindmelding. We called it like entering the hive mind. There was an article about this and I think like part of that is just that allows like the team to replenish. It's not that you I I was just on PTO in June. It's not just that you can take PTO and you come back to like 3x the amount of things to do.
[笑] 我觉得自己特别幸运的一点,是我们带出来的这个团队——既包括研究侧的那些 stakeholder,也包括我们研究产品管理团队内部的人。我觉得做这件事最神奇的地方在于,它不是一项个人运动。团队里有一种 radical ownership(极度当责)加上强协作的氛围。有时候确实感觉像在打一项高强度竞技,因为你面对的都是极关键的决策,用户侧、训练侧不断有新信息进来,你必须很快给出建议、判断和结论,而没有人能靠自己长期这么扛下去。所以真正帮到我的,是有一个非常了不起的团队,大家彼此照应——比如发布前一晚,哪怕自己不是那个模型的 DRI,也会陪着熬夜帮 DRI 一起审博客文章、改稿子、想更好的 demo,知道该在什么时候给对方当那只多出来的手。我觉得如果你把所有这些变化都扛在自己肩上,就很容易觉得孤立无援、觉得什么都得自己来。但我觉得 Anthropic 特别神奇的一点,就是我们有本事找出哪些地方可以互相搭一把,然后再多走一里路,做到某种意念同步。我们管它叫进入 hive mind(蜂巢思维),还有篇文章专门写过这个。我觉得这在很大程度上让团队能够回血。不是说……我六月刚休完 PTO,重点不是你能休假、然后回来面对三倍的活儿。
[1:19:11] Dianne
It's actually that you can take PTO and know the team can figure out the right things to do and that we individually can like watch out for each other. Um so I think that's a big part. I'm really lucky just personally um also my partner is really supportive um this is year six of me working in AI so Amazon and then anthropic and so he sees how much I just love the technology and what this can do and that really helps I think also um from like a personal perspective as well.
而是你休假的时候,能放心团队自己判断得出该做什么,我们每个人也都会互相看着点。所以我觉得这是很重要的一块。另外我个人也很幸运,我的另一半非常支持我。我做 AI 到今年是第六年了——先是 Amazon,然后是 Anthropic——所以他看得到我有多热爱这项技术、多相信它能做成什么,这一点从个人层面上也帮了我很多。
[1:19:44] Lenny
I love I love how many of these answers connect. So what I'm hearing here is just the having other people, working with other people, relying on other people, helping each other out when things get crazy. Uh which is a similar answer you had for just how to how to find the joy and and and fun in this work. Just get be inspired by other people, see what they're doing,
我特别喜欢这些回答之间的呼应。我听到的其实是:要有别人在身边,和别人一起做事,依靠别人,事情一疯狂就互相搭把手。这跟你刚才回答“怎么在这份工作里找到乐趣”几乎是同一个答案——被别人点燃,看看他们在做什么,
[1:20:05] Dianne
work together.
一起干活。
[1:20:06] Lenny
Yeah. And it's interesting when Fiona was on the podcast recently, she I was asking her just like what's changed in the world of software engineering and she pointed out it's a lot lonier now because now we're working with agents instead of other humans. Teams are smaller, people are have all these fleets they're talking to constantly. And so this is just a reminder of just the power of just actual other humans around you.
对。有意思的是,Fiona 最近上播客的时候,我问她软件工程这个领域到底变了什么,她说的是:现在孤独多了。因为我们现在是在跟 agent 一起工作,而不是跟别的人;团队变小了,每个人手里都带着一支 agent 舰队在不停对话。所以这正好提醒我们,身边真实的人到底有多重要。
[1:20:26] Dianne
We're we're asked to work and make decisions on really big things because you have more scale from the technology, right? And I think having individuals, having other folks more who can have some level of like mind meld with what you work on, how you approach maybe not exactly every detail, but what are the first principles? What are the assumptions you make then helps them uh you know back up for you or uh push your decision and sharpen your thinking. Um, so I think you know we really try to like I really try to look for that when like building the team, growing the team, hiring like is this person going to care about their own ego and building out a big org or are they going to care about contributing to anthropic and contributing to the like impact of the team and orienting towards folks who are like low ego team oriented. Um, I think that's, yeah, it it's a big part of I think the sustainability.
我们现在被要求去处理、去拍板的事情都非常大,因为技术给了你更大的规模,对吧?我觉得,身边有一些人、有更多同事能跟你在做的事上达成某种意念同步——不一定是每个细节都对齐,而是第一性原理是什么?你的前提假设是什么?——这样他们就能替你兜底,或者反过来推你的决定、把你的思考磨得更锋利。所以我们真的很努力去做这件事:我在搭团队、扩团队、招人的时候特别看这一点——这个人在乎的是自己的 ego、是把自己的组织做大,还是在乎给 Anthropic 做贡献、给团队的影响力做贡献?我会倾向于选那些低 ego、以团队为导向的人。我觉得这确实是可持续性里很大的一部分。
[1:21:30] Lenny
Yeah, just always a lot of it always just comes down back to culture and hiring and and I know I've heard a lot just the reason Anthropic is able to move so fast. I remember that moment when like something shipped every day of the month. There's like a calendar of launches and people were talking about how is this possible and what I heard a lot is just because everyone is so aligned around the mission and the values it allows people to make decisions really quickly before we get to our very exciting lightning round. Is there anything else Dan that you wanted to share? Anything else you wanted to touch on? Anything you want to maybe double down on of things we've talked about?
对,很多事情最后都会回到文化和招人上。我也听过很多说法,说 Anthropic 之所以能跑这么快……我还记得有那么一段时间,几乎每天都有东西发布,还有人做了一张发布日历,大家都在讨论这到底怎么做到的。我听到最多的答案就是:因为所有人在使命和价值观上高度对齐,这让大家能非常快地做决定。在进入我们非常精彩的 lightning round 之前,Dianne,还有什么你想分享的吗?还有什么想聊到的?或者我们聊过的哪些点你想再加重一下?
[1:22:04] Dianne
This was actually really fun because I feel like your questions actually sharpen some of my thinking around how the dots kind of connect. I'm I'm your real human claude over here. One thing that I really uh want to like convey or um have people take away is I think one in the ways of working, but also just two that like this is a this is a lot of like growth and change and having the joy in using this technology and like if you're feeling like in this moment you don't have as much of that feeling of initial joy, how do you find people who do uh if this is an area that that you're excited and like want to work on and I think developing skill sets replenishing skill sets in many ways of things like thinking from a first principles manner about what you solve I think fundamentally you didn't ask me this but there is this question in the community of do we still need PMS when the models are so capable when engineers are leaning in um I think the role of people who are user centric who go into the details of understanding what users are trying to accomplish bubbling that up in an actionable manner and doing the relentless work to do that like that to me is a core of a product person and I actually think we need more of that.
这次聊得真的很开心,因为我觉得你的问题反而帮我把一些思路串得更清楚了。我在这儿就是你的“真人版 Claude”。有件事我特别想传达、也希望大家能带走:一是工作方式上的;二是,这是一段充满成长和变化的时期,要在使用这项技术的过程中保有那份乐趣——如果你此刻觉得自己已经没有最初那种兴奋感了,那就去找那些还有这种感觉的人,前提是这确实是你兴奋、想投入的领域。另外就是培养技能、不断给技能补充新血,比如用第一性原理去想清楚你到底在解决什么问题。其实你没问我这个,但社区里一直有个疑问:模型都这么强了、工程师又这么主动上手,我们还需要 PM 吗?我觉得那些以用户为中心、愿意钻进细节去理解用户到底想完成什么、能把这些提炼成可执行的东西,并且愿意为此做那些不厌其烦的苦活的人——对我来说这才是产品人的内核,而且我认为我们需要更多这样的人。
[1:23:33] Dianne
I think we are becoming very technology layered driven and actually to make that impactful it's you have to go deep you have to be curious you have to be super hands-on and those are things that I think are also traits that have I think helped anthropic from a product development and model development perspective and as part of the culture and hopefully that's valuable for others as well.
我觉得我们现在正变得非常以技术层为驱动,而要让它真正产生影响,你必须钻得够深,必须有好奇心,必须极度亲力亲为。我觉得这些特质,也正是从产品开发和模型开发的角度上帮到 Anthropic 的东西,它们是这里文化的一部分,希望这对别人也同样有价值。
[1:23:59] Lenny
Amazing. What an inspiring way to end it. Oh man. Yeah. And this is I've been saying this too for a long time just now that building is easy the hard part part becomes as you said what should we build and is the thing we have built correct and good and worth leaning into and to me that's what PMs do and what PMs are good at.
太棒了。这个收尾太鼓舞人了。天哪。对,这话我自己也说了很久:现在“做东西”变容易了,难的部分变成了你刚说的那个——我们到底该做什么?我们已经做出来的这个东西对不对、好不好、值不值得继续押注?在我看来这正是 PM 在做的事,也是 PM 最擅长的事。
[1:24:17] Dianne
Yeah. Yeah. Yeah. And it's getting into the details of the user.
对对对。而且就是要钻进用户的细节里去。
[1:24:22] Lenny
Yeah. Empathy. Okay. Great. PMs are going to make it. Okay. PRD is not dead. [laughter] All kinds of all kinds of uh important lessons here. Uh Dan, with that we've reached our very exciting lightning round. I've got five questions for you. Are you ready?
对,共情。好,太好了。PM 有救了。好,PRD 没死。[笑] 这里面全是重要的收获啊。Dianne,说到这儿,我们就来到了非常精彩的 lightning round。我准备了五个问题,你准备好了吗?
[1:24:36] Dianne
Yep.
好的。
[1:24:37] Lenny
First question. What are two or three books that you find yourself recommending most to other people?
第一个问题。有哪两三本书是你最常推荐给别人的?
[1:24:43] Dianne
One personal one I really like how to raise an adult. So uh I'm a mom. I think a lot about what is the things that I want to instill in in in my kids. in that book is really helpful for describing we're not trying to raise children, we're trying to raise adults. So just the framing of what does that mean and what does it mean? What are the characteristics that we want to hone and like harness and foster in our kids? Um the other book that I uh was listening to on Audible recently is Incorable by Eric Reese. So the
有一本偏个人生活的,我很喜欢,叫 How to Raise an Adult(《如何养育成年人》)。我是一位妈妈,会经常想:我到底想在孩子身上种下些什么。这本书特别有帮助的一点是,它说我们要养的不是「孩子」,而是「未来的成年人」。光是这个框架就很有意思——这句话到底意味着什么?我们希望在孩子身上打磨、激发、培育的到底是哪些特质?另一本我最近在 Audible 上听的,是 Eric Ries 的 Incorruptible。
[1:25:22]
Incorruptible Incorruptible Yes. Yes.
Incorruptible。——Incorruptible,对,对。
[1:25:24] Lenny
Yeah. His recent podcast guest.
对,他最近还上过我的播客。
[1:25:26] Dianne
Um Yeah. And I I I just I think the question of how to build great companies is important. I personally just been most fascinated with how to keep great teams and great companies going further. And it was very interesting to just kind of see his framing and reframing of the question. Um I loved some of the examples around having metrics around culture. you if you can't if you only measure revenue and then that's kind of how you're going against but if you have other better metrics that's actually the way uh to to to sustain the the values you care about. I've been kind of trying to think about how to actually bring that to the team level of like how do we better articulate right our norms a lot of the things we talked about on the team. So I think [snorts] that's also a really good read.
对。我觉得「怎么建一家伟大的公司」是个很重要的问题,但我个人最着迷的其实是:怎么让一支很棒的团队、一家很棒的公司一直走下去。所以看他怎么给这个问题重新定义、重新提问,我觉得很有意思。我特别喜欢书里那些「给文化设指标」的例子——如果你只测营收,那你就只会朝营收去;但如果你有一些更好的指标,那才是真正能守住你在意的那些价值观的办法。我一直在想怎么把这套东西落到团队层面:我们怎么才能把团队里那些反复聊到的规范、默契,更清楚地表达出来。所以我觉得这本也很值得读。
[1:26:15] Lenny
There you go. Uh that'll be your next watch everyone as you're listening to this the Eric Greece episode. Yeah.
太好了。那各位听众,你们听完这期之后可以接着去听 Eric Ries 那一期。
[1:26:20] Dianne
Such a good episode. Yeah.
那期真的特别好。
[1:26:22] Lenny
And his book just came out. Incorruptible.
而且他的书刚出——Incorruptible。
[1:26:24] Dianne
Yes.
对。
[1:26:25] Lenny
And I think it was like a New York Times bestseller. Like it's actually doing incredibly well, which I was really happy to see.
我记得它还上了纽约时报畅销榜,卖得非常好,我看到挺替他高兴的。
[1:26:31] Dianne
Yeah, exactly.
对,就是这样。
[1:26:32]
Next question. Favorite recent movie or TV show you really enjoyed. Most people at Antropic don't have time to do what to watch things, but I'm curious if you have an answer. I would say um during uh some time off last month, I did get to like binge watch Fallout on Amazon Prime. So that was actually I kind of like um it's kind of uh Have you heard of it?
下一个问题。最近有没有特别喜欢的电影或剧集?我知道 Anthropic 的人大多没空看剧,不过还是好奇你有没有答案。——上个月我休了几天假,趁那阵子把 Amazon Prime 上的 Fallout(《辐射》)刷完了。我还挺喜欢的,它有点……你看过吗?
[1:26:56] Lenny
Yeah. Yeah, it's based on the video game.
看过看过,是改编自那款游戏的。
[1:26:58] Dianne
Yes, it's based on the video game. Uh I think it's a it was really um it's witty, it's humorous, it's also like super actionoriented. So highly recommend.
对,改编自那款游戏。我觉得它写得很机灵、很幽默,动作场面也特别足。强烈推荐。
[1:27:08] Lenny
Okay, next question. Do you have a favorite product you recently discovered that you really love?
好,下一个问题。最近有没有发现什么特别喜欢的产品?
[1:27:12] Dianne
I really do think like claw tag is very interesting in terms of a product experience. Um, we actually have like different versions of this uh within Anthropic and I I I think it's actually been really uh really really uh powerful tool.
我真心觉得 Claude Tag 作为一种产品体验非常有意思。我们在 Anthropic 内部其实有好几个不同版本,我觉得它真的是个特别特别强的工具。
[1:27:28] Lenny
Yeah, it feels like I think some people are like what's the big deal? The fact that everyone at Anthropic is like raving about it tells me something important is going on here. And I'm trying to actually get it working within my Slack community that I have for paid newsletter subscribers. How cool would that be?
对,我的感觉是——可能有人会说「这有什么了不起的」,但 Anthropic 上上下下都在为它疯狂打 call,这本身就说明这里面有点门道。我现在也在试着把它接进我给付费订阅读者开的那个 Slack 社区里。那得多酷啊。
[1:27:43] Dianne
Yeah.
是啊。
[1:27:43] Lenny
Yeah. I'm trying to figure out how it works when it's not a company when it's just a bunch of people that don't know each other and how that might work. But we're trying it out. Okay. Uh two more questions. Your favorite life motto that you find yourself often coming back to in work or in life. So I was actually raised by my grandparents uh for the first 10 10 years of my life and my parents were immigrant uh college and master students in the US.
对,我一直在琢磨:如果不是在一家公司里,而只是一群互不相识的人,这套东西还能不能跑得通、要怎么跑。反正我们正在试。好,还有两个问题。第一个:你在工作或生活里经常回想起来的人生信条是什么?
[1:28:08] Dianne
Oh wow. And um my grandfather always says, "No matter how far you go, there's always another level, [laughter] which uh um is I think um a really good way though, like a pretty uh intense way of describing uh his his life philosophy. But I go back to that whenever there's something new or unprecedented that we experience. And I think you know first half of this year there was definitely a lot of that like there was a lot of new things that we were learning. I was learning um so just feeling like there's always like another mountain another uh opportunity to
我头十年其实是爷爷奶奶带大的,我爸妈那会儿是在美国读本科和硕士的移民学生。我爷爷总说一句话:「不管你走多远,上面永远还有一层。」(笑)这话听着挺狠的,但我觉得它挺好地概括了他的人生哲学。每当我遇到全新的、前所未有的事情,我就会想起这句话。今年上半年这种时刻特别多,有太多新东西要学,我自己也在学。所以就是那种感觉——永远还有下一座山,永远还有下一个机会。
[1:28:50] Lenny
not good enough Dan we need to go better
「还不够好,Dianne,我们得做得更好。」
[1:28:53] Lenny
we need to go bigger. Uh makes me think about actually another Ben man line from his podcast episode that this is the most normal it's ever going to be. It's only going to get weirder and crazier.
「得做得更大。」这让我想起 Ben 在他那期播客里说的另一句话:现在已经是往后最正常的时候了,只会越来越怪、越来越疯。
[1:29:03] Dianne
Yeah. Yeah. No, we're good. Okay, final question. Uh, I was poking around at your LinkedIn. You were a high yield bond trader, JP Morgan Chase early in your career. Uh, you had like uh you have this like redacted uh hundred million dollar trading portfolio of some kind. Uh what did you learn from that time in your life that has stuck with you and or is there a crazy story from that period? It was four years of your life. I think I learned actually a lot that I uh apply here uh at at Anthropic and other uh jobs thereafter. Um so when I was at JP Morgan um the trading floor you could kind of envision like sort of Waffle Wall Street that's very different. Uh most traders I think are in front of a terminal. They're much more doing analyses uh on their computers. Um, but it's still very, I would say, like male-dominated. And so, uh, I was the only woman. I was the only, uh, um, person with like my background, uh, on the trading desk. And I learned that the it was a very good environment to kind of building one my sense of authentic self and two uh that even if I was the most junior person, even if I may look different, uh that the best ideas and having conviction in the best ideas uh irregardless of all of those other factors like is the most important thing. And so I think just bringing that sense of um how I show up more at work.
对对,没错。好,最后一个问题。我翻了你的 LinkedIn,你职业生涯早期在 JP Morgan Chase 做高收益债券交易员,还管过一个金额被打码的、上亿美元规模的交易组合。那段经历给你留下了什么、有什么一直跟着你到现在的收获?或者那段时间有没有什么疯狂的故事?那可是四年时间。我确实学到了很多,到现在在 Anthropic 和之前几份工作里都还在用。在 JP Morgan 的时候,交易大厅其实跟《华尔街》电影里那种想象很不一样——大部分交易员是坐在终端前面,更多是在电脑上做分析。但那个环境依然可以说是男性主导的。我是交易台上唯一的女性,也是唯一一个有我这种背景的人。我学到的是:那其实是个很好的环境,一来让我建立起真实的自我认知,二来让我明白,哪怕我是最资浅的那个、哪怕我看上去和别人不一样,最好的想法、以及对最好想法的笃定,才是最重要的,其他因素都不重要。所以我把这种状态带到了工作里。
[1:30:50] Dianne
Um I'm pretty vulnerable and authentic with my team. Uh I try to really make sure that regardless of people's levels or tenures, if they have a great idea, how to help them pursue that and to do also the same. Um so to like put the idea out there to actually um have conviction in it to do the follow through to do the like nitty-gritty work to make it happen. Um so those were all things that I learned from trading. Um and yeah I think applies to any any job in many ways.
我在团队里挺坦诚、也挺愿意示弱的。我会很努力地保证,不管一个人级别多高、资历多深,只要他有好想法,我就帮他把它推进下去,也希望他自己也这么做——把想法抛出来,真的对它有笃定,然后一路跟到底,把那些琐碎的脏活累活干完,让它落地。这些都是我从做交易那几年学到的。我觉得这些东西放在任何一份工作上都适用。
[1:31:20] Lenny
That is beautiful. Where can people find you online if they want to follow you and how can listeners be useful to you?
太棒了。如果听众想关注你,在网上哪里能找到你?他们能怎么帮到你?
[1:31:28] Dianne
I don't have a large presence on like uh social. Uh I think the best way to uh find my work uh my team's work is really uh the anthropic blog and when we're publishing new models, new product experiences I think in terms of uh useful uh for me I think the best thing number one is your feedback like we actually if if you thumbs up or thumbs down on any of our product surfaces if you contact your salesperson with feedback back about the model, it will make its way to me. Uh we actually with every like research model, I actually get pretty close into understanding favorability and feedback. Um so giving us that feedback, pushing Claude, telling us where it's falling down, um those help us make Claude better. Uh the other the other thing is like if you have folks in your network who seem like this type of profile of person that I just talked about I'm hiring the team is growing. We really will love just people who love this technology who are deeply curious first principles thinkers who are fearless in questioning assumptions um and who have like a tinkering hackery spirit.
我在社交平台上没什么存在感。想看我和我团队的工作,最好的渠道其实是 Anthropic 的博客,我们发新模型、新产品体验的时候都会在那儿。至于怎么帮我:第一,就是你们的反馈。我们真的会看——你在我们任何一个产品界面上点了赞或者点了踩,或者你把对模型的反馈告诉了对接你的销售,最后都会传到我这里。其实每一个研究模型,我都会相当深入地去看它的好评度和反馈。所以,把反馈给我们、去压榨 Claude、告诉我们它在哪儿掉链子,这些都能帮我们把 Claude 做得更好。第二件事是,如果你身边有符合我刚才描述的那类画像的人,我在招人,团队还在扩。我们特别想要的就是那种真心热爱这项技术、有极强好奇心、习惯第一性原理思考、敢于质疑各种前提假设,同时还带一点爱折腾、爱瞎鼓捣的黑客气质的人。
[1:32:46] Lenny
Wow dream job. So basically open open PM roles add anthropic on the research team.
哇,梦中情职。所以就是 Anthropic 研究团队这边有在招 PM。
[1:32:52] Dianne
Yes.
对。
[1:32:52] Lenny
And they apply I assume on the website the careers page.
我猜就在官网的 careers 页面投递?
[1:32:55] Dianne
Yes.
对。
[1:32:56] Lenny
Holy moly. All right. Here we go. Enjoy the flood of resumes you're about to receive.
天呐。行,那就这样。祝你享受接下来这波简历洪水。
[1:33:01] Dianne
Thank you. [laughter]
谢谢(笑)。
[1:33:03] Lenny
Uh Dan, thank you so much for being here.
Dianne,非常感谢你来上节目。
[1:33:06] Dianne
Thank you so much for having me. Thank you for um really helpful, thoughtprovoking questions um helping me even connect the dots on how how we work, how how this whole technology is coming together and being product people in it.
谢谢你邀请我。也谢谢你问了这些特别有价值、特别启发人的问题,帮我自己也把很多点串了起来——我们是怎么工作的、这整套技术是怎么拼到一起的,以及身处其中的产品人该是什么样。
[1:33:21] Lenny
I really appreciate that. But thank you Dan for real. Okay. Well, bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcast, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.
真的很感谢你这么说。不过说真的,该谢的是你,Dianne。好,那大家再见。非常感谢各位收听。如果这期对你有帮助,可以在 Apple Podcast、Spotify 或你常用的播客 App 上订阅本节目。也欢迎给我们打个分或者留条评论,这对其他听众发现这档播客真的很有帮助。所有往期节目和更多信息都可以在 lennyspodcast.com 上找到。我们下期再见。