ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.15 · 全文

The AI paradox: More automation, more humans, more work | Dan Shipper

频道: Lenny's Podcast
视频: https://www.youtube.com/watch?v=4D3hDmGhFhA
原文语言: en
统计: 共 50 轮 · Lenny 18 · Dan 25


[0:00] Lenny

The last time you were on this podcast, you had this hot take that people were sleeping on Claude Code. You were so unbelievably right. The premise of this episode is we're going to go through what else you predict will happen.

你上次上这档播客的时候,抛出过一个很有争议的观点——说大家都低估了 Claude Code。结果你说得太对了,准得离谱。所以这一期我们就想聊聊,你还预测会发生些什么。


[0:10] Dan

The AI job apocalypse is not really a thing. I am super super bullish on PMs and full-stack designers.

所谓的 AI 抢饭碗大灾难,其实根本不存在。我对 PM 和全栈设计师超级超级看好。


[0:18] Lenny

So, you guys are hiring doubled in people in the past year, which is not what people would have expected from a company that is so AI forward.

你们过去一年招人翻了一倍,团队规模扩了一倍。对一家这么 AI 优先的公司来说,这有点出乎大家意料。


[0:24] Dan

I'm simultaneously extremely AI pilled and [music] very bullish on humans. Automation is a lie. Every agent needs a human. We have so much automation, so much AI, and I also work way more.

我这个人很矛盾——一方面极度信仰 AI,另一方面又特别看好人。自动化是个谎言,每个 agent 背后都需要一个人。我们用了这么多自动化、这么多 AI,结果我反而工作得更拼了。


[0:34] Lenny

Creativity. It just feels like it's going to be more and more valuable to stand out from all the slop that people are shipping and launching constantly.

创意。我感觉,当大家都在不停地批量产出、批量发布那些粗制滥造的东西时,能脱颖而出会变得越来越值钱。


[0:40] Dan

What models do in general is they make yesterday's human competence cheap. And so, it becomes commoditized. It's not valuable anymore. What humans do is we go in there and we're like, "Yeah, we have all this frozen human competence from yesterday. How do I use this like make something new and interesting?" What are some predictions for how the way we work is going to change?

模型干的事,本质上就是把昨天还得靠人才能搞定的能力变得很便宜。于是这些能力就被商品化了,不再值钱。而人干的事是这样的——我们会进去,心想:好,我们手上有这么多昨天积累下来、被‘冻结’的人类能力,那我怎么用它做出点新鲜、有意思的东西来?那你对未来工作方式会怎么改变,有哪些预测?


[0:57] Dan

It's going to bifurcate in two main ways. One is everyone's going to have at least one agent that they talk to, that they can offload work to. Second is that most of the work that you do [music] is actually going to happen on your computer in an environment like Codex or Claude Co-work.

它会朝两个主要方向分化。第一,每个人至少都会有一个能对话、能把活儿甩给它的 agent。第二,你大部分的工作其实都会在电脑上、在像 Codex 或 Claude Cowork 这样的环境里完成。


[1:12] Lenny

What you're predicting here is the SaaS tools will run within Codex or Claude Code.

所以你在这里的预测是,那些 SaaS 工具会跑在 Codex 或 Claude Code 里面。


[1:17] Dan

I think the SaaS apocalypse is dumb. I would buy SaaS stocks right now. What agents do is increase the number of users of SaaS, not get rid of it.

但这并不是因为要「用更少的人做更多的事」,而是因为要做的事情实在太多了。


[1:24] Dan

A lot of people are moving to CLI and trying to work from the terminal. [music]

这里其实有一个特别有意思的悖论:你自动化得越多,需要的人反而越多。


[1:26] Lenny

We speed ran the CLI era. It was nice while it lasted, but I think CLIs are over. Today my guest is Dan Shipper, CEO and founder of Every. Dan and his team are building maybe the most AI forward startup out there. And as a result, are very much living in the future of how work is going to look as AI becomes a bigger and bigger part of our day-to-day. Everybody at their company, including every non-technical person, uses Codex and Co-work and Claude Code to get much of their work done. And this is why, way before [music] anybody else, Dan saw the rise of cloud code and what is now Cohere, which he predicted almost a year ago when he was on the podcast last [music] time. So, I asked Dan to come back on the podcast to share his current biggest predictions for how work is [music] going to change over the coming year for most people. We chatted about what work will look like at most companies at the end of this year, how the shape of the work we do will change, and who will do best in this coming future {slash} what you need to be working [music] on right now. Hint hint, product managers and designers are going to do very well. Dan makes a lot of bold predictions and many quite contrarian takes that I was not [music] expecting him to say, and we are going to revisit this conversation exactly a year from today to see how much he got right. Before we get into it, do not forget [music] to check out Lenny's Product Hunt dot com for a free year of the hottest and most well-crafted AI products in the world available exclusively to Lenny's newsletter subscribers. With that, I bring you Dan Shipper.

我们用最快的速度跑完了 CLI 时代。它存在的时候挺好用的,但我觉得 CLI 已经成为过去式了。今天我的嘉宾是 Dan Shipper,他是 Every 的 CEO 兼创始人。Dan 和他的团队打造的,可能是当今最 AI 优先的创业公司。也正因为如此,随着 AI 在我们日常工作中占据越来越大的比重,他们正实实在在地活在「未来的工作会是什么样子」当中。在他们公司,每个人——包括所有非技术岗的同事——都在用 Codex、Cowork 和 Claude Code 来完成自己大部分的工作。这也是为什么,远比其他所有人都更早,Dan 就预见到了 cloud code(Claude Code)的崛起,以及如今被称为 Cowork 的东西,这个他差不多一年前上播客时就预测到了。所以,我请 Dan 再次回到播客,分享他眼下对于「未来一年大多数人的工作将如何改变」最重大的几个预测。我们聊了今年年底大多数公司的工作会是什么样子,我们所做工作的形态将如何变化,以及在即将到来的这个未来里谁会过得最好——也就是你现在就该着手去做的事情。提示一下:产品经理和设计师会过得相当不错。Dan 抛出了很多大胆的预测,还有不少相当反主流的观点,有些是我完全没料到他会说的。而且我们打算正好在一年后的今天重新回顾这次对话,看看他到底说对了多少。在进入正题之前,别忘了去 Lenny's Product Hunt(lennysproducthunt.com)看看,那里有全世界最火、打磨得最精良的 AI 产品,免费送你一年使用权,且只对 Lenny's newsletter 的订阅者开放。说完这些,让我们有请 Dan Shipper。


[2:55] Dan

[music]

而我觉得其中特别难的一点是,那些走在前面的人,他们其实未必能说清自己到底在做什么,因为这一切太新了。所以我们在 Every 想做的,就是给大家提供一套语言和一套工具,让他们真正能把这些事做出来。


[2:56] Lenny

Dan, thank you so much for being here and welcome back to the podcast. Thanks for having me. Always a pleasure to be with you. The last time you were on this podcast, you had this kind of it was almost like an offhand hot take that people were sleeping on cloud code and in particular cloud code for non-engineering work, for just like fixing files, sorting your hard drive, just all these things that people hadn't thought about. Nobody was talking about this. This was a year ago. You were so unbelievably right about this. It's just like unreal what has happened since then. They built Cohere, which was this whole They built on this very specific idea using cloud code for non-technical work. A Codex is getting into this now. I imagine you've been seeing this. They're like leaning into this non-technical use of basically coding agents. I feel like this has also been a big part of Anthropic's success over the past year, just like how do non-technical people use this stuff? Uh so, you were just so go on this stuff. I I I even wrote a newsletter post building on this idea. I'm like, "Hey, this is you interesting. I should dig into this." I asked people, "How do you use cloud code for non-engineering work?" And I just had like so many examples and it's like my second most popular post. So, uh clearly you uh you you have a unique glimpse into where things are heading. So, the premise of this episode is we're going to go through uh what else you predict will happen in the future, how things will change for people building products. And I think it'd be helpful to start with giving people a brief glimpse into just how you operate and how your team operates that gives you this unique lens into where things are going. So, just give us a sense of how you how you work. Thank you. Um I I really appreciate the introduction. Um and yeah, I think what one of the things about predicting the future or or the way that we think about predicting the future at Every is that you what you don't want to do is prognosticate. What do you What you want to do instead is um is just live in it together. So, everybody at Every is an AI early adopter. We're almost 30 people now. I think when when we did our interview we were 15, so we've doubled in size in the in the last year. We're all early adopters and we have engineers, we have designers, we have writers, we have editors, we have um sales people, we have customer service people. And everybody has a little bit of that um whatever that thing is where you're just like, "Oh, I like to explore. I like to experiment. I'm very curious and I'm like super all in on AI." And what I what that does, I think, is it creates this like little pocket of the future where we're all living in it together and we get to be a little bit further ahead cuz at any other company there's like a mix of people. There's early adopters, there's like there's sort of like the middle of the pack people and there's people who are that like very anti. And another thing that happens, which is really cool, is we get to because of our role um you know, reviewing models and and being a little bit of a tastemaker in AI, we get access to stuff before it comes out. So, we get to beta test and alpha test and kind of help help steer the direction of where things are going a little bit, which is very, very cool. And so, when when I think about predicting the future, um it's actually when you create an environment like that, it's actually just about um noticing what's going on. Um and and I think what a core part of it, too, is writing about it. I think articulating what you're noticing, articulating the future, kind of brings it about in this way that um uh makes it real for you and your team and then anybody else who's like on the internet who's reading it. And so, the Claude Code thing, it was this is this very organic thing where for us, um we tried Claude Code when it came out. That's sort of our job. We we we try all the new stuff from all the new all the new model all we tried all the new stuff from the model companies. And at the time it was like a little bit early. But right around, I think like Sonnet 3.5 or Sonnet 3.7, we were testing that to do our vibe check on it. And we're like, "Holy [ __ ] This is crazy. This is like really You can They got rid of the code editor." And so, from that point on, we just basically we we run At this point now we run like six products software products internally. At that time we ran like maybe two or three. And from that point on, we just started shifting to a a world where everybody was No one was looking at the code. Everybody was you know, talking to their computer in English using Claude Code in the terminal. And so, I was able to see like, "Ooh, this is starting to happen." Um and then be because my job is a little bit to just like push and play with stuff, I was like, "I wonder if I could use this for like my writing. Like, how could I do that?" And then it just like starts to unfold and you're like, "Okay, this is not ready yet, but it's obviously useful for me. You know, my like one of the things that we talk about internally is what I call the reach test, which is like, do you just like, when you wake up in the morning, do you like reach for it organically? I love this combination of uh you're using the latest stuff, and I think this is, as you said, maybe an under underrated skill, you're you're good at uh being self-aware of here's what's weird and new and different and interesting. So, that's a really cool combination, partly cuz you have to write about it, and you write about it. So, I think that's like the perfect recipe for someone having a sense of where things are going. This episode is brought to you by our season's presenting sponser WorkOS. What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common? They are all powered by WorkOS. [music] If you're building a product for the enterprise, you've felt the pain of integrating single sign-on, SCIM, RBA, audit logs, and other features required by large companies. WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SaaS. Literally every startup that I'm an investor in that starts to expand upmarket [music] ends up working with WorkOS. And that's because they are the best. Whether you are a seed-stage startup trying to land your first enterprise customer or a unicorn expanding globally, WorkOS is the fastest path to becoming enterprise-ready and unblocking growth. It's essentially Stripe for enterprise features. Visit workos.com to get started or just hit up their Slack, where they have actual engineers waiting to answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience. Go to workos.com to make your app enterprise-ready today. So, the way that I'm going to structure this conversation, there's going to be basically three buckets of predictions. One is how the way we work is going to change in the coming years. Two is how what the shape of the work we're going to be doing is going to look like and change. And then three is who is going to be most successful in this future / what should you be doing and working on now to be successful in this future? Lenny, my only ask is we come on a year from now and then you score it. I want to score it. Okay, so this is a year from now. Okay. Okay. So, is this let's actually is [clears throat] this like your predictions for in a year this is what it's going to look like or this is like the emerging future?

[Lenny] Dan,太感谢你能来了,欢迎再次做客我们的播客。

[Dan] 谢谢你邀请我,能和你聊天永远是一种享受。

[Lenny] 上次你来这个播客的时候,你随口抛出了一个堪称「神预言」的观点——你说大家都低估了 Claude Code,尤其是低估了用 Claude Code 来做非工程类的工作,比如整理文件、给硬盘分类,所有这些大家根本没想到过的事情。当时没有人在谈论这件事。那已经是一年前的事了。结果你说得简直对到离谱。从那以后发生的一切真的让人难以置信。后来有人做了 Cohere(这里指的是基于这个想法做出来的东西),整个产品就是建立在「用 Claude Code 做非技术工作」这个非常具体的点子上。现在连 Codex 也开始往这个方向走了。我猜你应该一直都在关注这些。他们都在大力发掘这种「把编程 agent 用于非技术用途」的玩法。我感觉这其实也是过去一年里 Anthropic 成功的一大要素,就是「非技术人群怎么用这些东西」。所以呢,你当时真的太超前了。我自己甚至还顺着这个想法写了一篇 newsletter,我当时就想:「嘿,这个太有意思了,我应该深挖一下。」我去问大家:「你们都怎么用 Claude Code 来做非工程类的工作?」结果我收到了海量的例子,那篇文章成了我浏览量第二高的帖子。所以很明显,你对「事情会往哪走」有一种独到的洞察力。这一期节目的前提就是:我们要一起梳理一遍,你预测未来还会发生哪些事,对那些做产品的人来说,事情又会怎么变。我觉得,可以先让大家简单了解一下你本人是怎么工作的、你的团队是怎么运作的,正是这种工作方式给了你这种独特的视角去看清趋势走向。所以,先给我们讲讲你是怎么工作的吧。

[Dan] 谢谢。我真的很感激你这番介绍。是的,我觉得关于「预测未来」这件事——或者说我们在 Every 内部思考「预测未来」的方式——你最不该做的事就是「装神弄鬼地预言」。你真正该做的,是大家一起「活在那个未来里」。所以在 Every,每一个人都是 AI 的早期采用者。我们现在差不多有 30 个人了。我记得我们上次做访谈的时候是 15 个人,所以过去这一年我们的规模翻了一倍。我们全都是早期采用者,我们有工程师、有设计师、有写作者、有编辑,还有销售、有客服。而每一个人身上都多少带着那种特质——你懂的,就是那种「哦,我喜欢探索,我喜欢做实验,我超级好奇,而且我对 AI 是全身心投入」的劲头。我觉得,这么做的结果,是它创造出了一个「未来的小口袋」,我们所有人都一起活在这个口袋里,于是我们能比别人稍微领先一点。因为在其他任何公司,员工总是各种各样混在一起的:有早期采用者,有那种处于中间地带的人,还有那种极度抗拒的人。还有另一件事特别酷,就是因为我们的角色——你知道,我们要评测各种模型,算是 AI 圈里有点「定调者」味道的角色——所以我们能在很多东西正式发布之前就拿到使用权限。我们能做 beta 测试、alpha 测试,能稍微帮忙引导一下事情的走向,这非常非常酷。所以当我思考「预测未来」时,其实当你营造出那样一种环境之后,预测就只剩下「留意正在发生什么」这件事了。而我觉得,其中一个核心环节,是把它写下来。我觉得,把你所留意到的东西表达出来、把未来清晰地说出来,会以某种方式把那个未来「召唤」到现实里来,它会让那个未来对你、对你的团队,以及对互联网上任何读到它的人来说,都变得真实起来。所以 Claude Code 这件事,对我们来说是一件非常自然而然发生的事。Claude Code 刚出来的时候我们就试了,这某种程度上就是我们的本职工作——我们会去试所有模型公司出的所有新东西。当时它还有点太早期了。但大概就在 Sonnet 3.5 或者 Sonnet 3.7 那会儿,我们正在测它、给它做「氛围检验」(vibe check),然后我们就想:「我的天,这太疯狂了。这简直……他们把代码编辑器给干掉了。」从那一刻起,基本上——到现在为止,我们内部跑着大概六个软件产品;而那时候我们大概只跑着两三个——从那一刻起,我们就开始转向一个新世界:没有人再去看代码了。所有人都是在用 Claude Code、在终端里用英语跟自己的电脑对话。于是我就能看到:「哦,这事儿开始发生了。」然后呢,因为我的工作多少就是去「推一推、玩一玩这些东西」,我就想:「我好奇能不能把它用在我的写作上?我能怎么做到这一点?」然后它就这样一点点展开了,你会发现:「好吧,它现在还没完全准备好,但它对我来说显然是有用的。」你知道,我们内部经常聊到一个东西,我把它叫做「伸手测试」(reach test)——意思就是,当你早上醒来的时候,你会不会很自然地就「伸手去拿它」?

[Lenny] 我太喜欢这种组合了——你既在用最新的东西,而且我觉得这(正如你所说)可能是一种被低估的能力——你很擅长「自我觉察」,能意识到「这里有什么是奇怪的、新鲜的、不一样的、有意思的」。所以这真的是一种很酷的组合,部分原因是你必须去写它,而且你确实写了。所以我觉得,这就是让一个人对「事情会往哪走」拥有敏锐嗅觉的完美配方。

[Lenny] 本期节目由我们本季的冠名赞助商 WorkOS 带来。OpenAI、Anthropic、Cursor、Vercel、Replit、Sierra、Clay,以及其他数百家成功的公司,它们都有什么共同点?答案是:它们全都由 WorkOS 提供支持。[音乐] 如果你正在为企业级客户开发产品,那你一定体会过那种痛苦——要去集成单点登录(SSO)、SCIM、基于角色的访问控制(RBAC)、审计日志,以及大公司要求的其他各种功能。WorkOS 把这些「拖垮交易的拦路虎」变成了即插即用的 API,背后是一个专门为 B2B SaaS 打造的现代开发者平台。说真的,我投资的每一家初创公司,一旦开始向高端市场扩张 [音乐],最后都会用上 WorkOS。原因就是:他们是最好的。无论你是一家正在努力拿下第一个企业客户的种子轮初创公司,还是一家正在全球扩张的独角兽,WorkOS 都是「成为企业级就绪、并解除增长瓶颈」的最快路径。它本质上就是「企业级功能界的 Stripe」。访问 workos.com 即可上手,或者直接去他们的 Slack——那里有真正的工程师在等着回答你的问题。WorkOS 让你能更快地构建产品,提供令人愉悦的 API、详尽的文档和顺滑的开发者体验。今天就去 workos.com,让你的应用变得企业级就绪吧。

[Lenny] 那么,我打算这样来组织这次对话:基本上会有三大类预测。第一类是:在未来几年里,我们工作的「方式」会如何改变。第二类是:我们将要做的工作,其「形态」会是什么样子、会如何变化。第三类是:在这个未来里谁会最成功,以及为了在这个未来里取得成功,你现在应该做什么、应该在什么事情上发力?

[Dan] Lenny,我唯一的请求是:一年后我们再来一期,然后你来给这些预测打分。

[Lenny] 我想打分。好。

[Dan] 好,那这就是一年之后的事了。

[Lenny] 好。那我们其实先说清楚——[清嗓子] 这些到底是「你对一年后情况的预测,未来会变成这样」,还是说这是「正在浮现出来的未来」?


[9:57] Dan

like I don't I will probably say I don't have like an exact timeline. I think most of the stuff that I'm I'm going to talk about will be pretty apparent within a year, but it probably it may it may take longer than that. But I think it will it should within at least a year be like not obviously wrong. Like it it seem it could it should seem like it's moving in that direction to count. Okay. May of 2027, we will review your predictions. Right. Amazing.

(Dan)我想我大概会说,我没有一个特别精确的时间表。我觉得我接下来要讲的大部分内容,在一年之内就会变得相当明显,但也有可能会比这更久。不过我认为,至少在一年之内,它应该不会显得明显是错的。也就是说,它至少应该看起来是在朝那个方向发展的,这样就算数了。(Lenny)好的。那么 2027 年 5 月,我们来回顾你的这些预测。(Dan)对。(Lenny)太棒了。


[10:22] Lenny

Confirmed. Okay, I love this. Okay, so let's dive in. What are some predictions for how the way we work is going to change in the coming year? One of my favorite questions because I think if you look at the benchmarks, you're just looking at okay, like yeah, AI is going to just take all of our jobs basically, you know. I meter has this really cool benchmark where it's like it measures how long it can like the newest models can do tasks autonomously and it's like oh, it's like it can it what's it called? Oh, like mythos preview the like big anthropic model that everyone's like so worried about it can do tasks of 17 hours at 50% accuracy. It's like holy [ __ ] that's crazy. And I think it is real. It's true and and and the the progress like model progress is going up exponentially. And my experience and my feeling is that we will look back in a year and say we actually have a lot more work to do. Humans have a lot more work to do. Even as models get better at doing work and there's like a really interesting paradox there. And my prediction for the like how how work will Well, my my big prediction of how work will change or how you will be doing work in a year is it's going to bifurcate in this in two main ways, how you how you use agents. One is you're going to be doing I think like what we figured you would be doing like 5 years ago when we thought about how work with AI works, which is everyone's going to have at least in their company at least one agent that they talk to that uh can do work, that they can offload work to. And we'll talk about like what that looks like, but it's essentially like Open Claw. The second is that most of the work that you do is actually going to happen on your computer in an environment like Codex or Cloud Co-work that becomes the sort of operating system for it becomes the sort of operating system for how how you do all of your work, whether that's your email, the documents you create, like all that kind of stuff. It's going to be on that kind of a surface. It's that's becoming the the clear competitive landscape. Um so there's I want to go in order of those two. Um so the first one is you're going to have agents you delegate to, probably in Slack, but you know, anywhere. First thing that's interesting about that one is it's not clear what the architecture is going to be like for that. Um is everyone going to have an agent? Uh is every team going to have an agent? Is it going to be like just one agent? Is it like agent specializes? There's this parallel shadow org chart. And when Open Claw first came out, everyone internally at Every adopted it, and I was very convinced that it would be a everyone has their own agent. And there's like some real really interesting things about that world of, you know, a parallel a parallel org chart. Agents in that world sort of become little reflections of you, which is like really cool and really interesting. It's like if you ever Did you ever read The Golden Compass? Um it's like having a little daemon on your shoulder. You know, it's a little part of your soul. Um I I really think like that's sort of what it looked like was happening. Um and so I was very into personal agents. And I have completely flipped. And I I really think that uh the the model for now is going to be a super agent, like one agent for the entire company. And I you you're starting to see this in some companies. So like um Shopify very famously has one. Uh Ramp has one now. Um and and I think there's some like really interesting reasons for that. I actually still think that the personal agent thing is coming. But what we found is there's all this hype with Open Claw. Everyone's like, "I'm going to set it up. It's so cool." or whatever. And then everyone realizes it's like way too much work. This thing breaks all the time. I got to like fumble around with it. I got to be able to SSH into my server and like blah blah blah. And most people to do work at least just don't want to spend that time or can't. Um and the the like fundamental underlying thing that drives that is whether it's Open Claw or any other harness, in order for an AI agent to be useful right now, it really needs a human who cares about it. It really needs it like a human personal connection with someone who's like watching what it does and make sure that it's doing the right thing and that it's useful for people. And the minute you like sever that connection, so the minute someone's like, "Ah, like I don't I don't want to like maintain this like dumb Open Claw." is the minute the agent is like not really that useful anymore. And that's why it I think it has started to shift to a more uh one agent per company model because for now like the the the ideal is uh you you basically set up a forward deployed engineer or someone with that sort of profile who's responsible for making sure that that agent is working for the whole company. And then maybe you have some like some little team agents. Um and I think as the models get better at being more independent, that will like shift down and you will it'll be more likely that we'll have more personal agents cuz we don't have to [ __ ] around with all the internals. But um the model that I see working for us and for a lot of other companies, including the model companies, the model companies themselves are starting to see this, is when it comes to the sort of like async agents, it's really a you know, you have one agent at the top that's like doing sometimes it's everything, a lot of times it's um a particular kind of job that you've decided that everyone in the company needs an agent for like data requests. And uh and then I think it will start to it will start at the top at at the top and then it sort of starts to trickle down where you make it more specialized agents and teams and and all that kind of stuff. And the mechanism is agents need people who care about them. That is so interesting that point about you need to like garden your agent because there's context you have to keep adding to it. There's like it breaks as you said and it's just like once it's just too much work, you're like, okay, forget this thing. I'm going to go back to CodeX or Cloud or something like that. Exactly.

[Lenny] 确认好了。好,我太喜欢这个话题了。好,那我们就开始吧。对于未来一年我们工作方式将如何改变,你有哪些预测?这是我最喜欢的问题之一,因为我觉得如果你去看那些 benchmark,你看到的基本就是:是啊,AI 差不多要把我们所有的工作都抢走了,你懂的。[Dan] METR 有个特别酷的 benchmark,它衡量的是最新的模型能自主完成多长时间的任务,然后你会发现——哦,它能做……那个叫什么来着?哦对,就是那个大家都特别担心的、Anthropic 的大模型 mythos preview(指 Claude 的某个预览版大模型),它能以 50% 的准确率完成需要 17 小时的任务。你会想,我靠,这太疯狂了。而且我觉得这是真的,确实如此,模型的进步是在指数级地往上走。但我的经验和我的感觉是,一年之后我们回头看,会发现其实我们要做的工作反而更多了。人类要做的工作更多了。哪怕模型越来越擅长干活,这里面有一个特别有意思的悖论。那么我对工作将如何……嗯,我对工作会如何改变、或者说一年后你会怎么干活的重大预测是:它会朝两个主要方向分化,也就是你使用 agent 的方式会分成两类。第一类,是我觉得就像我们五年前设想用 AI 工作时所预期的那样——每个人在公司里至少会有一个可以对话的 agent,这个 agent 能干活,你可以把工作甩给它。我们待会儿会聊这具体长什么样,但它本质上就像 Open Claw 那种东西。第二类,是你做的大部分工作其实会发生在你自己的电脑上,在一个像 Codex 或者 Claude Co-work(即 Cloud Co-work,自动字幕误识)这样的环境里,这个环境会变成某种操作系统——会变成你完成所有工作的那种操作系统,不管是你的 email、你创建的文档,所有这类东西,都会在那样一个界面上完成。它正在成为那种界面。这就是正在形成的、清晰的竞争格局。嗯,所以这两类我想按顺序来讲。第一类,就是你会有一些可以委派任务的 agent,可能是在 Slack 里,但其实在哪儿都行。关于第一类,第一个有意思的点是:它的架构到底会是什么样的,目前还不清楚。是每个人都有一个 agent 吗?还是每个团队有一个 agent?还是说就只有一个 agent?还是说 agent 会专门化分工?这就形成了一张平行的、影子般的组织架构图。当 Open Claw 刚出来的时候,我们 Every 公司内部所有人都采用了它,而我当时非常确信,未来会是每个人都有自己专属的 agent。那个世界里——你懂的,一张平行的组织架构图——有一些特别特别有意思的东西。在那个世界里,agent 某种程度上变成了你自身的小小投影,这点真的特别酷、特别有意思。就好像……你读过《黄金罗盘》(The Golden Compass)吗?嗯,就像你肩膀上有一个小小的精灵(daemon),它是你灵魂的一部分。嗯,我真的觉得当时看起来就是这么回事。所以我当时特别迷恋个人 agent 这个方向。但我现在彻底反转了。我现在真的认为,目前阶段更合适的模式会是一个超级 agent,就是整个公司共用一个 agent。你已经开始在一些公司里看到这种做法了。比如说,Shopify 就有一个非常出名的例子。Ramp 现在也有了一个。嗯,我觉得这背后有一些特别有意思的原因。我其实仍然认为个人 agent 这个趋势终会到来。但我们发现的是,Open Claw 引发了这么多炒作,每个人都说:「我要把它配置起来,它太酷了」之类的。然后大家就意识到——这玩意儿工作量也太大了。它老是出故障。我得不停地折腾它。我得能 SSH 进我自己的服务器,然后各种 blah blah blah。而大多数人,至少为了干活而言,就是不愿意、或者没法花那个时间。嗯,而驱动这一切的、最根本的底层原因是:不管是 Open Claw 还是任何其他 harness,一个 AI agent 想要在现阶段真正有用,它真的需要有一个在乎它的人。它真的需要一种类似人与人之间的私人连接——需要有个人盯着它在做什么,确保它做的是对的事、确保它对大家是有用的。而一旦你切断了这种连接,也就是说一旦有人开始想:「啊,我不想再去维护这个破 Open Claw 了」,那一刻,这个 agent 基本上就变得不那么有用了。这就是为什么我觉得现在已经开始转向「一家公司一个 agent」的模式,因为就目前而言,理想做法是:你基本上安排一个 forward deployed engineer(前置部署工程师)或者具备这类画像的人,由他负责确保那个 agent 能为整个公司正常运转。然后你可能再配一些小的团队级 agent。嗯,我觉得随着模型越来越擅长独立工作,这个重心会往下移,我们也会更有可能拥有更多的个人 agent,因为到那时我们就不必再去折腾所有那些内部细节了。但目前我看到的、对我们以及很多其他公司——包括那些模型公司本身——都行之有效的模式(这些模型公司自己也开始意识到这一点)是:在涉及那种异步(async)agent 的场景里,真正有效的其实是——你懂的——你在最顶端有一个 agent,它有时候什么都干,但很多时候是干某一类特定的活,就是那种你已经决定全公司每个人都需要一个 agent 来处理的事,比如数据请求(data requests)。然后我觉得它会从顶端开始——从最顶端开始,然后逐渐向下渗透,在这个过程中你再把它拆分成更专门化的 agent、分配到各团队,诸如此类。而背后的机制就是:agent 需要在乎它们的人。[Lenny] 你说的这点太有意思了——你得像打理花园一样去「照料」你的 agent,因为你得不断给它补充 context。而且就像你说的,它会出故障,到最后你会觉得:一旦工作量太大,你就会想,算了,不管它了,我还是回去用 Codex 或者 Claude(Cloud,自动字幕误识)之类的东西吧。[Dan] 完全正确。


[16:39] Lenny

Okay, cool. So this is a cool opportunity. So the idea that So what you're predicting here is uh companies will have this super agent that everyone can talk to. As you said, a Shopify's got River, I think it's called. What's the Ramp one called? I can't remember. Okay. It's probably got a fun name. Okay. So uh So that's the prediction. Okay. That's that's the first prediction.

好的,酷。所以这是个很棒的机会。所以这个想法是……所以你在这里预测的是,呃,各家公司都会有这么一个超级 agent,每个人都可以跟它对话。就像你说的,Shopify 有个叫 River 的(我记得是这个名字)。Ramp 那个叫什么来着?我记不清了。好吧,它大概也有个挺有意思的名字。好的。所以呢,这就是那个预测。好。这就是第一个预测。


[16:57] Dan

That's the first prediction.

这是第一个预测。


[16:58] Dan

um we will start with agents at the top that uh that are more general and are used by more people in the company and then it will start to kind of grow down as the as people get more used to these use cases, they get more specialized and um agents become uh less uh uh less fiddly. Like they just work better. And is this mostly going to be in Slack, do you predict? For work? Yeah, it seems to make sense. I think people I people love having the green bubbles on Open Claw. Um like it I'm sorry, the the blue bubbles on Open Claw. Like if you can use it with your iPhone, but I think there's this little thing in people's heads where they really like to keep their personal and work agents separate. Mhm. And um I think there's a whole there's a whole territory Our our COO Brandon Gall calls this um computer errands. There's like a this whole territory of using personal agents for your computer errands. It's like order my groceries or whatever, and it's like there's so much of that that I think this is going to it's going to be huge for, but um I focus we focus mostly on the work stuff. Um and I think that's going to happen mostly in Slack. Sweet. Go Slack. Should we Do you want to talk about the uh the other work surface?

嗯,一开始 agent 会出现在最上层——那种更通用、公司里更多人都会用的 agent。然后它会慢慢往下渗透:随着大家越来越熟悉这些用法,agent 会变得更专门化,也不再那么"难伺候",就是说它们会越用越顺。 (Lenny)那你预测这主要会发生在 Slack 里吗?工作场景下? 对,这挺说得通的。我觉得大家其实挺喜欢 Open Claw 上的绿气泡的……抱歉,是蓝气泡。就是说,如果你能在 iPhone 上用它当然好,但我觉得人们心里有个小执念——他们特别想把个人 agent 和工作 agent 分开。我们的 COO Brandon Gall 把这块叫做"电脑跑腿"(computer errands):用个人 agent 帮你处理电脑上的杂事,比如帮我下单买菜之类的。这类需求太多了,我觉得它会非常巨大。但我们主要聚焦在工作这一块,而工作这块大部分会发生在 Slack 里。 (Lenny)太好了,Slack 加油。那我们要不要聊聊另一个工作界面?


[18:12] Dan

Absolutely. Codex co-work. Okay. This is the last Let's do it. I'm so excited about this one. I think it's the coolest thing. So, basically what happened was Anthropic realized at some point that with Claude Code, if you put an agent on your computer and it runs on your computer, it has everything it has access to everything that you have access to. It uses the terminal, so it has like basically superpowered access to it. And not only that, it really these agents really understand how to use the terminal cuz there's so much content online about about that. And it it created this like superpowerful coding paradigm, which is um you know, Anthropic was really doing it first. OpenAI for a while was I I I in my opinion like very very behind on this, and then in my opinion has surpassed them recently. It's really interesting. Um but they were very early on this. Um when people were still thinking about coding agents or coding models as being really pair programmers, they were among the first to be like, "No." And do it successfully. Like there were people before them like Devin who I think had had a big had the big like cloud environment and and Open AI tried this too, but um but the the real adoption seems to have happened when you uh put it on your computer. So, they figured that out. And then I think they figured out um along with their community that once you have a coding agent on your computer that can build anything, it's actually really good for any kind of work you want to do. And people started just hacking Claude code essentially to do all of their work. So, Anthropic then built co-work um which is you know, a little bit of a nicer wrapping around Claude code, but it's fun- fundamentally the same thing. And then I think you know, I think Open AI made a couple of different bets, but their main bet on a programming agent was the the the the earlier version of the Codex were like very technical and they were like super smart, but they were like a little bit autistic. Like it was a little hard to they they didn't quite get what you meant. They get ex- they got exactly what you said. And I think maybe like three or four months ago around the time that they launched uh 5.3, they started to move in this direction of, "Oh, no, we get it. Like it's um this model is fast. It's like really good for general purpose knowledge work type tasks." And then they launched the Codex desktop app. And I think the Codex desktop app takes If you look at all the lessons that like Anthropic learned, they went from Claude code to co-work. And you can kind of see that in the tabs on the on the Anthropic desktop app UI. I think Open AI was just like, "We We see where this is going. Like let's just skip to that." And so, I think Codex right now this is a horse race. Like it they're going to have different positions. Um but I I think Codex right now is my daily driver. I like spend all all my time in it basically. I flip the card every once in a while, but I think they're getting the paradigm right and it's clear to me that whoever is in the lead cuz I again I think it'll change. Whoever is in the lead, it feels very obvious to me that all of the work that you do is going to be in one of those surfaces where uh for example, when I'm writing a document Codex has a browser in uh in the app. It has an in-app browser. And when I'm writing a document, I just go into one of my uh one of my Codex threads which I have one thread for every project. And I just open the in-app browser. I go to the document. I usually do it in Proof which is this um online mark markdown editor that I built. And then I just have Codex running and watching me in Proof. And Codex can see what I'm doing. I can see what Codex is doing. It's all kind of in one place which is the an extension of the same thing that made Claude code work really well originally. And I basically feel like I have this parallel work buddy that not only can it like respond and write in the document, but then it can go do research. It can go it can use my computer to basically do anything that I can do on my computer. And that's like incredibly powerful. Um and I do this with everything. Like I've been in I've been in inbox zero for like 10 days straight now which if you know me is crazy. I'm never like this. And that's because I literally just have Codex gather all my emails with Cora which is our email agent. And then um it it renders a little page uh and I I think I showed you this at the Anthra- at the Anthropic event. It renders a little page and I just like monologue into it and just talk at each email. I'm like, "Okay, go go research this. Oh, here's a question from our lawyers. Can you go like collect all of the you know, documents from the last like four years and then put them into a report and send them?" And it just does it. And so all the stuff that I would procrastinate on, I don't really procrastinate on anymore. And so I feel like there's this For a long time, we thought I thought, too, that the optimal experience of AI was going to be take AI and put it in a browser. And I think the reverse is actually starting to happen and be like really, really valuable in a way that I did not expect, which is take the AI agent that you use all the time on your computer and put a browser in it so it can see everything you're doing. And that is just like a magical combination that I think will be is very uncommon now. You can't even do this in cloud in cloud code um because they they don't let you browse external websites inside of cloud code. So it's very uncommon now, but I think it will be super common in a year. This is more profound than it may even sound. What I'm hearing is instead of AI being baked into SaaS tools, what you're predicting here is uh you will the SaaS tools will run within Codex or Claude code. That That is That is one uh really important uh second-order effect of this is um Okay, so yeah, like I'm I'm using Proof or or really any website, maybe PostHog or whatever. And I'm doing it inside of my agent. And the agent has access to the website, so it has access to everything that I have access to. And it has access to my whole computer. When I run the agent on that website, I'm using my tokens. I'm not using the the vendor's tokens. I'm not using the app's tokens. And so it puts SaaS back in its place where yeah, you want to make it friendly for an agent. And everyone's got a CLI now. Um you want to make the HTML uh really uh really usable. You want to make sure that what anything that happens in the CLI shows up for the user immediately. All that kind of stuff. There are a lot of issues to to deal with. But um once you do that, you actually don't really need to think about having a an AI surface that's primarily going to be the thing that users use in the sense that you don't need to build an agent naturally into your product. I think you can and there's there's another really interesting bifurcation of this that that that we should talk about which is that having two agents is better than one. Um but I think for now there's this really cool thing where uh with proof for example, uh anyone who uses it, I don't pay for tokens because they're just bringing they bring their AI to the to proof. And so it changes what you build as a SaaS company uh and you build it now for both humans and agents to use at the same time and it changes your margins back to well, I don't really have to pay for tokens anymore cuz the user is going to bring AI. So I think this is a huge deal. So what you're describing here is uh more and more work that we do, more and more professional work is it just going to happen within Codex or Cloud Code. Uh how where does Cursor fit into this? Is that one of the is is there potential there?

当然。Codex 的 co-work。好,这是最后一个了,来吧——我对这个特别兴奋,我觉得它是最酷的东西。 基本上是这样:Anthropic 在某个时间点意识到,对 Claude Code 来说,如果你把一个 agent 放在自己的电脑上、让它在你电脑里跑,那它就拥有了你拥有的一切权限。它用的是 terminal,所以基本上获得了一种超级强大的访问能力。而且不只是这样——这些 agent 是真的很懂怎么用 terminal,因为网上关于 terminal 的内容实在太多了。这就造就了一种超强的编程范式。Anthropic 确实是最早做这个的;OpenAI 有一阵子在我看来非常非常落后,然后最近又反超了,这点挺有意思的。但 Anthropic 在这上面起步很早。当大家还把 coding agent 或 coding model 当成"结对编程伙伴"来看的时候,他们是最早站出来说"不,不是这样",而且还做成了。在他们之前也有像 Devin 这样的,我记得 Devin 有那种很大的云端环境,OpenAI 当时也试过,但真正引爆采用率的,似乎还是当你把它放进自己的电脑里之后。 他们想明白了这一点。然后我觉得他们和社区一起又想明白了:一旦你有一个能在电脑上构建任何东西的 coding agent,它其实非常适合干任何你想干的活儿。于是大家就开始拿 Claude Code 各种"魔改",用它来处理自己所有的工作。所以 Anthropic 后来就做了 co-work——你可以理解成给 Claude Code 套了一层更好看的壳,但本质上是一回事。 再说 OpenAI,他们下了好几个不同的注,但在编程 agent 上的主要押注是这样的:早期版本的 Codex 非常技术、非常聪明,但有点"轴"——你想表达什么它不太能 get 到,它只会严格照你说的字面去做。大概三四个月前,差不多在他们发布 5.3 的时候,他们开始转向另一个方向:"哦,我们懂了,这个模型很快,特别适合通用的知识工作类任务。"然后他们推出了 Codex 桌面应用。 我觉得 Codex 桌面应用是把 Anthropic 学到的所有经验都拿过来了。Anthropic 是从 Claude Code 一步步走到 co-work 的,你在 Anthropic 桌面应用 UI 的那些标签页里能看出这个演进过程。而 OpenAI 大概就是想:"我们看清这玩意儿的走向了,干脆直接跳到终点。"所以现在 Codex 和它们是一场赛马,大家会有不同的定位。但 Codex 现在是我的主力工具(daily driver),我基本上所有时间都泡在里面。偶尔我也会换着用,但我觉得他们把范式做对了。而且对我来说很明显,不管现在领先的是谁——我还是觉得格局会变——但不管谁领先,未来你做的所有工作都会发生在这类界面里。 举个例子,我写文档的时候,Codex 应用里自带一个浏览器,一个 in-app browser。我写文档时,就进到我的某个 Codex 线程里(我给每个项目都开一个线程),打开内置浏览器,进到那个文档。我一般用 Proof——这是我自己做的一个在线 markdown 编辑器。然后我就让 Codex 一直跑着、看着我在 Proof 里干活。Codex 能看到我在做什么,我也能看到 Codex 在做什么,全都在一个地方——这其实就是当初让 Claude Code 那么好用的那套东西的延伸。我基本上感觉自己有了一个并肩工作的搭档,它不仅能在文档里回应、写东西,还能去做调研,能用我的电脑去干任何我能在电脑上干的事。这太强大了。 我什么事都这么干。我已经连续 inbox zero 十天了——认识我的人都知道这有多疯狂,我从来不是这样的人。原因就是我直接让 Codex 用 Cora(我们的邮件 agent)把我所有邮件收集起来,然后它会渲染出一个小页面——我在 Anthropic 那个活动上给你看过——它渲染出一个小页面,我就对着它"碎碎念",对着每封邮件说话:"好,去查一下这个;哦,这是律师的一个问题,你能不能把过去四年的所有相关文件都收集起来、整理成一份报告发出去?"它就真的去做了。所以那些我以前会一直拖延的事,现在我基本不拖了。 有很长一段时间,我以为——我也曾这么想——AI 最理想的形态是把 AI 塞进浏览器里。但我现在觉得正好反过来的事情正在发生,而且价值大到出乎我意料:是把你一直在电脑上用的那个 AI agent,往里面塞一个浏览器,让它能看到你在做的一切。这种组合简直是魔法。现在这还很少见,你甚至没法在 Claude Code 里这么做,因为它不让你在 Claude Code 里浏览外部网站。所以现在很罕见,但我觉得一年后会非常普遍。 (Lenny)这件事比听上去还要深远。我听到的意思是:与其说 AI 被烤进 SaaS 工具里,你这里预测的是——SaaS 工具会跑在 Codex 或 Claude Code 里面。 对,这是其中一个很重要的二阶效应。比如我用 Proof,或者其实任何网站,可能是 PostHog 之类的,我都是在我的 agent 里面用它。agent 能访问这个网站,所以它拥有我拥有的一切权限,同时还能访问我的整台电脑。当我在那个网站上跑 agent 时,用的是我自己的 token,不是厂商的 token,不是那个 app 的 token。这就把 SaaS 重新放回它该在的位置:你要做的是让产品对 agent 友好。现在大家都有 CLI 了,你要让 HTML 真正好用,要确保 CLI 里发生的任何事都能立刻显示给用户,诸如此类。要处理的问题很多。但一旦你做到了,你其实就不太需要去专门搭一个"主要给用户用的 AI 界面"了,也就是说你不必非得把一个 agent 硬塞进你的产品里。我觉得你当然可以塞,而且这里还有另一个特别有意思的分叉值得聊——就是"两个 agent 比一个好"。但眼下有个很酷的现象:拿 Proof 举例,任何人用它,我都不用为 token 付费,因为他们是自带 AI 来用 Proof 的。所以这改变了你作为一家 SaaS 公司要构建的东西——你现在要同时为人类和 agent 来构建产品。而这也把你的利润率改回去了:我不太需要再为 token 付费了,因为用户会自带 AI。所以我觉得这是件大事。 (Lenny)所以你描述的是:我们做的越来越多的工作、越来越多的专业工作,都会发生在 Codex 或 Claude Code 里面。那 Cursor 在这里面处在什么位置?那块有没有潜力?


[25:52] Dan

That's a good question. I think that Cursor sees a lot of the same stuff. And they're and in some ways they have some of the same stuff but it's better. Like I think that Cursor's cloud implementation is better than either it or when AI's or Anthropic's and is more advanced. And I think that Cursor has at least so far more distinctly chosen a lane. Like they're more distinctly choosing to be a for programmers. And that may limit how far they get in here. Like I think the definition of programmer is expanding enough that they'll have a big market, but I don't know that they're going to jump into like okay, use this to make a slide deck or whatever. But it is really clear that every model company is starting to realize how important it is to have a harness to get the most out of the um the model. And so where the where all the platforms are moving is to a world where you're not just doing prompt and response when you call the the model at on on the open AI platform, the Anthropic platform, you are they are literally like running the model on a computer that that is in the cloud that they run and then giving you the result out of it. And they know that they in order to get the best results of the model they need to offer that and so you see, you know, Anthropic's got cloud managed agents. Um open AI does not have a have a response yet, but I assume that that's going to happen and now Cursor uh was just essentially acquired by SpaceX. It's not like a full acquisition, but it's close. So I think people are starting to realize like I can't just do the like model part of it. I have to have this like harness above it and I think the ultimate form of that harness is like I can do any kind of knowledge work. Cursor itself is feels like one of the things that it's going to be a hard decision for them whether to stay just for coders or not. So people building products that aren't open AI or Anthropic, if this proves to be true, the prediction here is they're going to be using your product over time inside of one of these agents. Uh is there something you would do if you were one of those companies to prepare for that future? I would I would just prepare for that. So like, you know, for for example, um every more classic piece of productivity software, whether it's Slack or uh Word docs or PowerPoints or whatever, it's really mostly meant for a human to use. Um and now people are doing CLIs, so it's like meant for uh an agent to use independently of a human. And we're moving into this new paradigm, I think, where the human and the agent are on the same piece of work together and they're both doing things and you need to have I need to have visibility into what the agent is doing. The agent has to have visibility into what I'm doing. We have to go back and forth in this sort of like seamless way. And the kind of software that you make for that is going to be very different. So, for example, um like there's a lot of stuff that Proof doesn't have. I don't have to have a lot of the like Word document kind of like formatting or page breaks or like, you know, making tables or whatever cuz the agent just does it. I don't need to worry about that. It can do all the formatting for me. So, you can make the products a lot simpler and faster to start than the legacy products are. And then there's all these other affordances that you need to start to have because the way agents interact with software is very different. So, for example, agents can do a lot at once. They can just do like a billion different things to your document or your slide deck or your code base or whatever. And how you display that to the user is going to be very different than the way you might display a human being concurrent in your document and doing stuff. You need um you need like approval. You need a sort of inbox that sort of summarizes, here's all the stuff that's going to happen or has happened. You need um you need logs and the ability to roll it back real quick. So, there's all those kinds of considerations that um that change the actual product. And then the underlying UX of it or the underlying infrastructure you need is different, too, because, you know, agents can make a billion requests in like 3 seconds. So, how are you going to deal with that, right? Um this is exactly why, you know, GitHub is having problems right now cuz because the the number of people using GitHub has is skyrocketing exponentially and it's really just people's agencies in GitHub. So, I I think it's a this whole new world that is just starting you're just starting to see like a little peek of it. But, there's so many cool things about it. So, for example, in Proof and some of our other products, too, uh when someone has a problem, they don't email support. Their agent sends a bug report. And an agent bug report is way better than a human bug report. Um it has like, here's exactly what I did, here's the exact repro steps, here's like Proof is open source, so here's what I think is going in the code base. And then we just get that, it becomes a GitHub issue, and then we can just like send off an agent to fix it. And um you can't do that with everything, but it's so much better. And you can see the like the glimmers of this this very fast like closed loop between I ran into something, a paper cut, a little feature I want, a little bug, and my agent just goes off and talks to the company agent, and then the company agent just goes and fixes it. That I think is incredibly cool. So, is there a part of this that you a lot of people are moving to a CLI and trying to work from the terminal? Is part of this prediction that people shift away from that and back to actually you you acts with agents kind of running alongside them? CLIs are over. Um we we speed ran the CLI uh era. It was nice while it lasted, but I think it's pretty it's pretty clear It's not that CLI Sorry. It's not that CLIs are going to completely go away. Obviously, they've been around for the last like 30 years or 40 years or 50 years or whatever. They will continue to be around. And I think there is this moment when cloud code was like so popular and uh or or when when cloud code was really starting to gain in popularity, that people were like the the thing that's working is the fact that it's a CLI, and I don't think that's what it is. And when you move into an actual UI for this, you start to realize um we made GUIs for a reason. And it's just nicer to be in a GUI. And you can get all the same benefits inside inside of a GUI, especially for non-programmer work, but I would I would estimate that definitely the majority of the technical people inside of every are not using CLIs anymore as their main work surface. I think a lot of programmers are still flipping into it every once in a while, but it's more or less they're using Codex, cloud code, cursor, um that kind of thing. Awesome. Okay. I I I would I definitely wanted to make that part clear. So, coming back to kind of the the big picture of the prediction here, there's kind of these two modes of work that you're anticipating. One is this kind of super agent within a company that you chat with through Slack most likely that can go off and do work and answer questions. And then there's on your computer running Codex or Cloud Code. And within that all the work that you normally do kind of on your computer is now going to be living within Codex or Cloud Code or maybe some third-party that emerges that we're not not even aware of yet. Yes, and you're going to use apps inside of the internal browser of those of those tools. Wow, okay. Like listening to you talk about it, it's like it may not feel as profound as it is cuz this is a big change to how we work. We don't currently have an AI that we talk to regularly in Slack and we also don't work currently mostly in Codex or Cloud Code. So, this is actually a pretty massive shift. I think so. Is there anything else along these lines before we get into our next prediction? Well, a few things. I'm definitely not an agent maximalist. Like I really think we're going to have a lot of different agents that we use. Seems pretty clear to me. And I really do think that two agents are better than one. So, what's a good example? When I have Codex interact with another agent, it can give so much more context about me and what I want than I would be able to type. And it can go back and forth talking about things that would take a long time for me to express directly to an agent that you get this like speed up effect when you assume that your users are you are using Codex or Cloud Code or Co-work as their as their basic way they access your app. And a really simple example, we have this um hosted open claw product which we we had it we had it on wait list. We actually had to deposit because we started taking people off the wait list and open claw is just a very hard agent harness to to make work. It's like it's moving so incredibly fast. Uh and if you're like a platform for it, it just it's like it when things break you can't fix it. It's very hard. Um but one of the things that we learned in that process is if you're let's say you're building an agent product um or or a new any new software experience, what you would assume, let's say to set up an agent, is you need to build like a little like web interface or a little um slack workflow that ask people about, "Okay, like who are you and um what are you going to use this for and like what's your what's your ideal, you know, dream outcome?" Or what like whatever the things you are that you would put on an on onboarding checklist. If instead you you just you just make a hard line of we are only going to service users who use Codex or Co-work. Um what happens is you just paste something into you just paste a prompt into Codex or Co-work. It goes and talks to the app and the app can be either just a regular server or it or it can be its own agent. And Codex has so much information about you that it can just give it here's all the stuff I've been working on with Dan. Um here's all the ways that you know, he might he might want to use this app and then bring it back to me and it's this very custom experience. And also for a technical product like an agent, when something goes wrong, I can just tell Codex, "Go fix it." And Codex will go talk to the app and figure out what's going on for me. And so I think the whole paradigm starts to change when you assume that everyone's got an agent and those agents are talking to other agents in this like really magical and important way. There's a couple more things I want to touch on before we get started cuz there's like so much to talk about. Uh one of you made this point about SAS tools not using like you can use tokens from the uh model companies basically when using a SAS tool. Talk a bit more about that cuz that may change the business model for SAS companies in the future. That feels like a big deal. Well, I think it actually may uh save their margins. Because right now all the all these companies are rushing to like add a agent to their offering and thinking, "Oh, the agent is going to be the main way that I that people interact with me." And I think that uh and that costs tokens, obviously. And I actually think once I have once I have Codex or Co-work as my main work surface, I still want to use SAS. So, this is another good prediction. I would buy SAS stocks right now. Um I would I think the SAS apocalypse is done and SAS stocks will be up majorly in the next couple years. Not Not investment advice, but you know, I would buy SAS stocks. Um So, So, uh so I think it saves your margin because now what you're what the way that you're thinking that is not I have to build AI into this. It's It's more like I have to make a piece of software that humans and AI want to collaborate on together. And that's hard, but it's once you build it, it's a lot cheaper than assuming everyone's spending tokens. And um I it's I think it's a I think it's a good business. And And part of the reason I'm so bullish on SAS is A, everybody internally here is uh like I said, we've all got agents and we're all using Codex and whatever and we still pay for a ton of SAS and our SAS spend is up year over year. And we're not like vibe coding every single like little thing, you know? And I think that what agents do is increase the number of users of SAS. Not get rid of it. And so I think SAS companies are going to see like an insane spike in the amount of demand that they have because there's going to be tons of agents using these products at like a very high volume. And like I said, that's a huge infrastructure challenge. There's a There's a lot of like interesting pricing challenges, but uh it it it makes me very bullish on SAS. I love that if anything else comes out of this conversation, Dan Shipper, SAS is the future of AI. This

问得好。我觉得 Cursor 看到的东西跟大家差不多,某些方面他们有同样的东西、但做得更好。比如我觉得 Cursor 的云端实现(cloud implementation)比 OpenAI 或 Anthropic 的都要好、更先进。而且至少到目前为止,Cursor 更明确地选定了一条赛道——他们更明确地选择做"给程序员用的"。这可能会限制他们在这个方向上能走多远。我觉得"程序员"的定义正在扩大到足够大,他们会有一个很大的市场,但我不确定他们会不会跳进"用这个来做一份幻灯片"之类的场景。 但有件事非常清楚:每一家模型公司都开始意识到,要把模型的能力榨到极致,外面套一层 harness 有多重要。所以所有平台都在往这个方向走:当你在 OpenAI 平台或 Anthropic 平台上调用模型时,已经不只是"prompt 进、response 出"了——他们实际上是在自己运行的云端电脑上把模型跑起来,再把结果给你。他们知道,要拿到模型最好的结果,就得提供这种能力。所以你看 Anthropic 有 cloud managed agents;OpenAI 目前还没有对应的产品,但我猜这一定会来;而 Cursor 现在基本上被 SpaceX 收购了——不算完全收购,但很接近。所以大家开始意识到,光做"模型"这部分是不行的,必须在它上面有这一层 harness,而这层 harness 的终极形态就是"我能做任何一种知识工作"。Cursor 自己感觉会面临一个艰难的抉择:要不要继续只做程序员的工具。 (Lenny)那对于那些不是 OpenAI 或 Anthropic、自己在做产品的公司来说,如果这个预测成立,意思就是:随着时间推移,他们的产品会被人放在这类 agent 里面来用。如果你是这些公司之一,你会做点什么来为这个未来做准备吗? 我会就照着这个去准备。比如那些更经典的生产力软件——不管是 Slack、Word 文档、PowerPoint 还是别的——它们其实主要是为人类使用而设计的。现在大家开始做 CLI,那就是为 agent 独立于人类去使用而设计的。而我觉得我们正在进入一个新范式:人和 agent 一起处理同一份工作,两边都在动手,所以我需要能看到 agent 在干什么,agent 也得能看到我在干什么,我们要以一种无缝的方式来回配合。为这种场景做的软件会很不一样。 举例来说,Proof 里很多东西是没有的。我不需要 Word 文档那种排版、分页符、做表格之类的,因为 agent 直接就帮我做了,我不用操心,它能帮我搞定所有格式。所以你能把产品做得比那些老牌产品简单得多、上手快得多。然后又有一堆新的"affordance"(功能可供性)是你必须开始具备的,因为 agent 跟软件交互的方式很不一样。比如 agent 一次能干很多事,它可以对你的文档、幻灯片、代码库一口气做上亿种改动。那你怎么把这些展示给用户,就会跟展示"一个人类正在文档里同时操作"非常不同。你需要审批机制;你需要一种类似收件箱的东西,把"接下来要发生的全部改动、或者已经发生的全部改动"做个汇总;你需要日志,还要能快速回滚。所以有一大堆这类考量会改变产品本身。 而且底层的 UX、底层需要的基础设施也不一样,因为 agent 能在 3 秒内发出上亿个请求。那你打算怎么应对?这正是为什么 GitHub 现在遇到麻烦了——用 GitHub 的"人数"在指数级飙升,而其实那大部分是大家的 agent 在用 GitHub。所以这是一个全新的世界,你现在只是刚刚瞥见它的一角而已。但它有太多很酷的点了。 比如在 Proof 还有我们其他一些产品里,当有人遇到问题时,他们不发邮件找客服,而是他们的 agent 发来一份 bug 报告。而 agent 写的 bug 报告比人写的好太多了:它会写清楚"我具体做了什么、确切的复现步骤是什么";Proof 是开源的,它还会说"我觉得代码库里出问题的地方在这儿"。我们拿到它,它就变成一个 GitHub issue,然后我们直接派一个 agent 去修。这事不是什么场景都能做到,但能做到的时候真是好太多了。你已经能看到这种非常快的闭环初现端倪:我撞上了一个小麻烦、一个想要的小功能、一个小 bug,我的 agent 就跑去跟那家公司的 agent 对话,然后公司的 agent 就去把它修好。我觉得这太酷了。 (Lenny)那这里面有没有一部分是……现在很多人都在转向 CLI、想从 terminal 里干活。你这个预测里是不是也包含:大家会从那种方式撤回来,回到"agent 在你身边并行运行"的形态? CLI 时代结束了。我们把 CLI 时代"速通"了一遍。它存在的那段时间挺美好的,但我觉得已经挺清楚了……不是说——抱歉——不是说 CLI 会彻底消失。显然它已经存在了三四十年、五十年了,以后也会继续存在。我想说的是:当 Claude Code 特别火、或者说当它人气真正开始起来的时候,大家会说"它之所以好用,关键就在于它是个 CLI"。我不觉得是这样。当你为它做一个真正的 UI 时,你会开始意识到——我们当初造 GUI 是有原因的,待在 GUI 里就是更舒服。你能在 GUI 里拿到所有同样的好处,尤其是非编程类的工作。我可以估计:Every 内部绝大多数技术人员,现在已经不再把 CLI 当作主要工作界面了。我想很多程序员还是会时不时切进去用一下,但大体上他们用的是 Codex、Claude Code、Cursor 这类东西。 (Lenny)太好了。我特别想把这一点讲清楚。那回到这个预测的大图景,你预期会有两种工作模式。一种是公司内部的那个"超级 agent",你很可能通过 Slack 跟它聊天,它能去干活、回答问题。另一种是在你电脑上跑 Codex 或 Claude Code,而你平时在电脑上做的所有工作,现在都会活在 Codex、Claude Code、或者某个我们现在还没意识到、将来会冒出来的第三方里。 对,而且你会在这些工具的内置浏览器里使用各种 app。 (Lenny)哇,好。听你这么讲,它可能听上去没有它实际那么深远,因为这对我们的工作方式是个巨大的改变。我们现在并没有一个在 Slack 里经常对话的 AI,我们现在大部分工作也不在 Codex 或 Claude Code 里。所以这其实是一次相当巨大的转变。 我也这么觉得。 (Lenny)在进入下一个预测之前,关于这条线还有别的要补充的吗? 有几点。我绝对不是一个"agent 极大化主义者"。我真心觉得我们会用很多很多不同的 agent,这在我看来很清楚。而且我真的认为两个 agent 比一个好。举个好例子?当我让 Codex 去跟另一个 agent 交互时,它能提供关于我、关于我想要什么的、远比我自己打字能给出的多得多的 context。它能就一些事情来回沟通——而那些事如果让我直接对一个 agent 表达清楚会花很长时间。当你假设你的用户是用 Codex、Claude Code 或 co-work 作为访问你 app 的基本方式时,你就能拿到这种"提速"效果。 一个很简单的例子:我们有个托管版的 Open Claw 产品,本来是放在等候名单(wait list)上的,我们甚至得收押金,因为我们开始从等候名单里放人进来——Open Claw 实在是个非常难做稳的 agent harness,它迭代快得离谱,如果你是给它做平台的一方,东西一坏你根本修不了,非常难。但我们在这个过程里学到的一点是:假设你在做一个 agent 产品、或者任何新的软件体验,按常规思路,要"设置一个 agent",你会以为你得做一个小小的 web 界面、或者一个 Slack 工作流,去问用户"你是谁、你打算拿它来干嘛、你理想中的梦想结果是什么"——也就是你会放进新手引导清单里的那些东西。但如果你反过来,划一条硬线:"我们只服务用 Codex 或 co-work 的用户",那会发生什么呢?用户只要往 Codex 或 co-work 里粘贴一段 prompt,它就会去跟你的 app 对话(这个 app 可以只是一台普通服务器,也可以本身就是个 agent)。而 Codex 关于你的信息太多了,它可以直接把"这是我一直在跟 Dan 一起做的所有事、这是他可能会怎么用这个 app"全都给过去,再把结果带回给我,这就是一种非常定制化的体验。而且对于一个像 agent 这样的技术产品,出问题的时候我可以直接告诉 Codex"去修好它",Codex 就会去跟那个 app 对话、帮我搞清楚出了什么状况。所以当你假设"每个人都有一个 agent、而这些 agent 又在以这种很神奇、很重要的方式互相对话"时,整个范式就开始变了。 (Lenny)在我们继续之前我还想多碰几个点,因为可聊的实在太多了。你刚才讲到 SaaS 工具——用 SaaS 工具时基本上可以用模型公司那边的 token。能再多讲讲这个吗?因为这可能会改变 SaaS 公司未来的商业模式,感觉是件大事。 嗯,我觉得这其实可能会拯救他们的利润率。因为现在所有这些公司都在抢着往自己产品里"加一个 agent",心想"agent 会成为大家跟我交互的主要方式",而这显然要烧 token。但我其实觉得,一旦我把 Codex 或 co-work 当成我的主力工作界面,我仍然想用 SaaS。所以这又是一个不错的预测:我现在就会买 SaaS 股票。我觉得"SaaS 末日"已经结束了,未来几年 SaaS 股票会大涨。不构成投资建议哈,但我会买 SaaS 股票。 所以我觉得这能救你的利润率,因为现在你的思路不再是"我必须把 AI 做进产品里",而更像是"我必须做出一个人和 AI 都愿意一起在上面协作的软件"。这很难,但一旦做出来,它比"假设每个人都在烧 token"要便宜得多。我觉得这是门好生意。 我之所以这么看好 SaaS,部分原因是:第一,我们内部所有人,像我说的,人手一个 agent、都在用 Codex 之类的,但我们仍然付费买了一大堆 SaaS,而且我们的 SaaS 支出还在逐年上涨。我们也不是每个小东西都去 vibe coding。我觉得 agent 做的事是增加 SaaS 的用户数量,而不是干掉它。所以我觉得 SaaS 公司会迎来需求量的疯狂飙升,因为会有海量的 agent 在以非常高的频率使用这些产品。像我说的,这是个巨大的基础设施挑战,也有一堆有意思的定价难题,但这让我对 SaaS 非常看好。 (Lenny)我太喜欢了——如果这场对话能留下点什么,那就是:Dan Shipper 说,SaaS 才是 AI 的未来。这个……


[38:53] Lenny

[laughter]

(笑)


[38:55] Lenny

You to be SAS. #sendtweet I I love just Yeah, this is quite contrarian and the other interesting piece is that the fact that you guys are hiring that you doubled in people in the past year, which is not what people would have expected from a company that is so AI forward. Talk about what your experience there of just Okay, we still actually need humans. Automation is a lie. Um in the sense that every time you automate something, in order to make sure the automation is working well, you need a human on top of it like making sure that it's working well. And so um you know, I wrote this piece a couple years ago called the allocation about the allocation economy like the idea that the way that humans are going to work with AI is going to is going to be like like being a manager. And the thing that you have to remember about managers is like managers actually spend a lot of time working. Most managers are not like on the beach. They're like checking in with their employees all the time and and and and trying to figure out Okay, how do we make this work good? How do we make it better? How's it doing? How's this person doing? All that kind of stuff. And I think there's there's some differences between being a human manager and being a model manager, but um fundamentally it it still requires a lot of time and attention. And I think that we kind of miss that in the model discourse. And one of the reasons is benchmarks make it look like AI is more autonomous than it is. And by autonomy, I mean something specific by autonomy, and I'm going to try to express this. It's like a little hard to express, but I learned this for myself because I've been feeling this paradox a little bit. I've been feeling the like we have so much automation, so much AI, and I also work way more.

你要成为 SaaS。#发推。我太喜欢了。是啊,这观点相当反共识。另一个有意思的点是:你们其实在招人,过去一年人数翻了一倍——这可不是大家会从一家如此 AI 先行的公司身上预期到的。聊聊你在这方面的体会吧,就是那种"我们其实仍然需要人"。 (Dan)自动化是个谎言。我的意思是:每当你把一件事自动化,为了确保自动化运转良好,你就需要一个人在上面盯着,确保它跑得对。所以——我几年前写过一篇文章叫《配置经济》(the allocation economy),核心想法是:人类未来跟 AI 协作的方式会很像当一个管理者。而关于管理者你必须记住一点:管理者其实花了大量时间在工作上。大多数管理者并不是躺在沙滩上的,他们一直在跟员工对接、不停地琢磨"我们怎么把这事做好、怎么做得更好、现在进展如何、这个人状态怎么样",诸如此类。我觉得当"人类管理者"和当"模型管理者"之间是有一些区别的,但本质上它仍然需要投入大量的时间和注意力。我觉得我们在关于模型的讨论里把这一点给漏掉了。其中一个原因是:benchmark 会让 AI 看起来比它实际上更自主(autonomous)。我说的"自主"是有特定含义的,我来试着表达一下——这有点难说清楚,但我是从自己身上学到的,因为我一直在切身感受这个悖论。我一直在感受这种:我们有这么多自动化、这么多 AI,可我反而工作得多得多。


[40:36] Lenny

[snorts]

(噗嗤一笑)


[40:36] Dan

And I think part of the paradox or part of the paradox started to like resolve for me a little bit when I made my own benchmark. So, I made this senior It's called a the senior engineer benchmark, and it's like, "How good is AI versus a human engineer?" And the way that I built it is again, have this app proof. I just vibe coded it on the side and I like while running the rest of every. And when we launched it, because it was completely vibe coded, it just started going down and I couldn't fix it. And it was very embarrassing. I had a lot of egg on my face. And like the product worked. We We tested it internally. We had a lot of beta testers, but like the day after launch, it was like just every like 10 minutes the servers would go down and people were looking at me and I'd be like, "I don't know what's going on." Like, "Codex, fix it." And Codex was like, "I don't know what's going on." Um or really Codex was like, uh "I do know what's what's going on. I fixed it." And then it it would cause four other errors, and then you're just going around in a circle and I wasn't sleeping, and I I I vibe coded so hard I got bursitis on my elbow. So, uh that's a There's a life lesson in there. Vibe coder elbow.

我觉得这个悖论开始对我稍微解开了一点,是在我自己做了一个 benchmark 之后。我做了一个……它叫"资深工程师 benchmark"(senior engineer benchmark),就是衡量"AI 跟一个人类工程师比到底有多强"。我做它的方式还是老一套:用 Proof 这个 app,我就在一边随手 vibe coding 把它搞出来,同时还在运营 Every 的其余一切。结果我们一上线,因为它完全是 vibe coding 出来的,它就开始不停宕机,而我修不好,特别尴尬,把我糗得够呛。产品本身是能用的,我们内部测过,也有很多 beta 测试者,但上线第二天,差不多每隔 10 分钟服务器就挂一次,大家都看着我,我只能说"我也不知道咋回事"。我就跟 Codex 说"修一下",Codex 说"我也不知道咋回事"。或者更确切地说,Codex 会说"我知道咋回事了,我修好了",然后它一修就又冒出另外四个错误,于是你就这么原地打转。我那阵子根本没睡,vibe coding 太猛了,胳膊肘都得了滑囊炎(bursitis)。所以——这里头有个人生教训:vibe coder 网球肘。


[41:44]

[laughter]

(笑)


[41:45] Dan

Um so, anyway, I got a I got actually two different senior engineers to fix it independently. So, I have two different rewrites of the code base that um tells me how they did it, right? And so, what I get to do is when we get new models I just give the new model a prompt. I say like, "This is vibe coded slop. If you wanted to rewrite it from first principles, how would you write it? Go do it. And all the models until GPT 5.5 got like a 30 out of 100. And senior like a human senior engineer gets like high 80s, low 90s out of 100. So, there's a lot to go. And then I tried GPT 5.5 and it got like a 62. And mind you, the 60 the 60 score was um GPT 5.5 using an Opus 4.7 plan. Opus 4.7 plans are very good. GPT 5.5 is the only model though that has the sense of agency and confidence to just like rip out old code and just like actually rewrite from first principles. Other coding models, they kind of like try they like end up papering over the edges around the edges and they're like, "Oh, this is a big job. Like I'll just do a little patch." And you're like, "No, I like specifically told you not to." So, GPT 5.5 there's like a 30-point bump in the score. 60 out of 100. It's like very it's very clear that in a year or less, it's going to be senior engineer level, right? And that gives you a certain picture in your mind, especially based on how I named the benchmark, which I think a lot of benchmarks do. And I can tell you that when we get to that point, I will be very it will be very easy for me to change the benchmark to zero out the current model. So, that gets a zero out of 100. And so, for example, uh it seems like there's no skill or no thought into the prompt, which is this is vibe code is slop like fix it from first principles, but actually it took me a a while to get to a prompt that didn't give away the answer, but uh uh but got the model to reveal what it's capable of. And the original prompt I gave it was the original prompt that I gave it when uh when I was trying to fix the issue in production was going down, which is like I'd woken up I'd I'd woken up in the morning and I was like, "Okay, we had four or five reported issues yesterday. I want you to go through all the issues and then come to like a make a plan for how to resolve all of them and go do it, right?" And every coding model on the market, and I I am I'm pretty sure this Here's a prediction. I'm pretty sure every coding model on the market will still do this in a year. Every coding model on the market will take that instruction seriously. And if I tell it, "Here's a bunch of issues, go fix it." They will just go try to fix the issues. What a actual human senior engineer does is they go look at the code base and they're like, "This is a piece of [ __ ] This guy doesn't know what he's doing."

总之呢,我后来找了两位资深工程师,让他们各自独立地把这套东西修好。所以我手上就有了两份对这套代码库的重写版本,能让我看清他们各自是怎么做的。这样一来我能玩的花样就是:每次有新模型出来,我就给它一个 prompt,说:「这是一坨 vibe coding 出来的烂代码。如果让你从第一性原理出发重写,你会怎么写?去做吧。」结果在 GPT 5.5 之前的所有模型,得分都只有 100 分里的 30 分左右。而一个真正的人类资深工程师能拿到 80 多分、90 出头。所以差距还很大。然后我试了 GPT 5.5,它拿到了大概 62 分。而且要说明一下,这个 60 分是 GPT 5.5 用了一份 Opus 4.7 的方案才达到的——Opus 4.7 出的方案非常好。但 GPT 5.5 是唯一一个有那种主动性和自信、敢于直接把旧代码连根拔起、真正从第一性原理重写的模型。其他写代码的模型呢,它们往往只是修修补补,最后在边边角角打打补丁,还会说:「哦,这是个大工程,我先打个小补丁吧。」你只能无语:「不,我明明专门告诉你别这么干。」所以 GPT 5.5 在分数上一下子跳了 30 分,到了 100 分里的 60 分。可以很清楚地看出,一年甚至更短时间内,它就会达到资深工程师的水平。这会在你脑子里勾勒出一幅特定的图景,尤其是考虑到我给这个 benchmark 起的名字——我觉得很多 benchmark 都有这个问题。但我可以告诉你,等真到了那一天,我会非常轻松地把这个 benchmark 改一改,让现在的模型重新归零。也就是说,让它重新拿个 0 分。举个例子,表面上看这个 prompt 好像没什么技巧、没什么心思——就是「这是 vibe coding 的烂代码,从第一性原理修好它」——但其实我花了挺久才打磨出一个既不直接泄露答案、又能让模型把自己真正的能力暴露出来的 prompt。我最初给它的那个 prompt,是当初线上出问题、系统快崩了的时候我用的那个:我早上醒来,心想「好,昨天我们收到了四五个报告的问题,我要你把所有这些问题过一遍,然后做一个解决所有问题的计划,然后去执行」。而市面上每一个写代码的模型——我相当肯定,这里给个预测——我相当肯定一年后市面上每一个写代码的模型还是会这么干。它们都会把这条指令当真照办。如果我跟它说「这里有一堆问题,去修」,它们就真的会一头扎进去逐个修这些问题。而一个真正的人类资深工程师会做的是:他去翻一遍代码库,然后说「这是一坨翔,写这个的人根本不知道自己在干什么」。


[44:45]

[laughter]

(笑)


[44:45] Dan

And then then they say, "We're going to have to like actually rewrite a lot of this and it's going to be hard and risky. I know you don't want to hear that, but like we're going to have to do that." And if you asked the model, "Hey, like should we do that?" It'll it'll probably it it'll probably get there, but it's not going to do it on its own. Um and it and there's a lot of incentives pushing against it doing that. And even if it does that, there's a there's always a higher frame for us to go. And so I think it's it's really important uh when when we think about benchmark progress to think about it from that perspective, which is benchmarks rise on problems that we've framed that we can articulate, that we can score. And there's a lot of work that's human work that uh it it can't be scored until you write it down, but the act of thinking to prompt it or write it down um is uh is something that you can't measure, but like kind of means that even if the benchmarks get saturated, it doesn't mean the same thing as we you totally replace all senior engineers. And it's I think it's why even though the models are getting better at automation, I still hire engineers. I am so excited to tell you about this season's supporting sponsor, Vanta. Vanta helps over 15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers. Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been. Every security leader that I talked to is feeling the increasing weight of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do without dated solutions. Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA. This helps companies get compliant fast and stay compliant. More than ever before, trust has the power to make or break your business. Learn more at vanta.com/lenny. And as a listener of this podcast, you get $1,000 off Vanta. That's vanta.com/lenny. One thing I mentioned recently on the podcast, I heard that speaking of the code that you have of like humans writing code uh data labeling companies are buying code that was written before 2021, 2022, before AI became a thing is like very valuable data.

然后他会说:「我们恐怕得把这里面很大一部分真正重写一遍,这会很难、也有风险。我知道你不爱听这个,但我们必须这么做。」如果你去问模型「嘿,我们该不该这么做?」,它大概也能想到这一层,但它不会主动去做。而且有很多因素在阻止它这么做。就算它真这么做了,我们也总还能往上再拔高一层框架去看问题。所以我觉得,在思考 benchmark 上的进展时,从这个角度去看真的很重要:benchmark 之所以能涨分,是因为这些问题是我们已经框定好的、能清晰表述的、能打分的。而还有大量工作是属于人类的工作——在你把它写下来之前,它根本没法被打分。但「想到去给它提 prompt、或者把它写下来」这个动作本身,是你没法度量的,可它某种意义上意味着:哪怕 benchmark 被刷满了,也不等于「我们彻底替换掉了所有资深工程师」是一回事。我想这也正是为什么,尽管模型在自动化上越来越强,我还是在招工程师。我特别想跟大家介绍本季的赞助商 Vanta。Vanta 帮助超过 15,000 家公司——比如 Cursor、Ramp、Duolingo、Snowflake、Atlassian——赢得并证明客户对他们的信任。多亏了 AI,团队构建和交付产品的速度比以往任何时候都快。但与此同时,引入到你产品和业务中的风险也比以往任何时候都高。我接触到的每一位安全负责人,都越来越深切地感受到保护自己组织、业务,以及客户数据的重担。因为一切都变得太快,他们只能不停地被动应付,靠猜来排优先级,还得将就着用过时的方案。Vanta 把合规与风险管理自动化,覆盖 35 多种安全与隐私框架,包括 SOC 2、ISO 27001 和 HIPAA。这能帮企业快速合规、并持续保持合规。在今天,信任比以往任何时候都更能决定一家公司的成败。想了解更多请访问 vanta.com/lenny。作为本播客的听众,你还能享受 1,000 美元的优惠,记住是 vanta.com/lenny。我最近在播客里提到过一件事——说到你这种由人类写的代码——我听说数据标注公司正在收购 2021、2022 年之前、也就是 AI 还没火起来之前写的代码,这些被当作非常有价值的数据。


[47:15] Dan

Original human code. Yeah, exactly. That's exactly right. And it's so interesting that that's exactly the kind of code used to build this model. Well, what's interesting? So, I want to I want to clarify there. So, I did not have a human write the code all by hand. Because I actually think that that's sort of it feels silly to me. Like, I don't really care because I know if if an engineer is not using AI, like I'm not going to work with them. I don't really care. It's like it's sort of like am I going to race a human against a car? Like, I probably wouldn't do that. But, um I would race a human in a car versus another human in a car and say which one's better. And in this case, what the the way the benchmark is structured is yeah, like these human engineers used AI, but they used it in a way that I could not cuz I didn't understand it and I didn't have time and I didn't really want to like go in and try to understand the code base, to be honest. And I think that's a really important thing when we think about benchmarks is AI is a broadly distributed technology that any human can use and when we are benchmarking against humans, AI against humans, we're actually really always talking about one human using AI versus another human using AI cuz AI doesn't use itself. It it may be able to in this like slightly somewhat recursive way, but there's in any real use case, there's always a human like pretty close to it making sure that it's working. Okay, I'm going to try to wrap up our first bucket. There's so much to talk about. I've made a little list of things that I think people should do based on your predictions to be successful. We'll talk about this at the end, too, but just a few things. One is start using Codex or cloud code more and more for the work you're doing and especially the browser, use tools inside of it. Two is allow your allow agents to be to use your products. If you're really going to SAS tool, make it easy for agents to be a user, essentially. Three is start thinking about some Slack bot that you can work with, like try out tools. Like I know Slack has their own Slack bot that I think is really good, too, and I haven't played with it, but people really like it. So, look for I guess a tool that could become the AI agent within your company. Buy SAS stock ASAP.

纯人类写的代码。对,完全正确,就是这样。而且特别有意思的是,恰恰就是这类代码被用来训练出了这些模型。不过有意思的是……我想在这里澄清一下:我并不是让人类纯手工把代码全写出来。因为我其实觉得那样有点傻。我并不在乎这个——因为我很清楚,如果一个工程师不用 AI,那我根本不会跟他合作。我不在乎。这就好比,我会让一个人去跟一辆车赛跑吗?我大概不会。但我会让一个人开车去跟另一个人开车比,看谁更强。在这个例子里,这个 benchmark 的设计方式就是这样:是的,这些人类工程师用了 AI,但他们用 AI 的方式是我做不到的——因为我看不懂、也没时间,老实说我也不太想真的钻进去搞懂那套代码库。我觉得这是思考 benchmark 时一个非常重要的点:AI 是一种被广泛普及的技术,任何人都能用。当我们拿人来做对照、拿 AI 对人来比时,其实我们谈的永远都是「一个用 AI 的人」对「另一个用 AI 的人」,因为 AI 不会自己用自己。它或许能以某种略带递归的方式做到一点,但在任何真实的使用场景里,背后总有一个离它很近的人在确保它正常运作。好,我想试着把我们第一个话题收个尾。要聊的实在太多了。我列了一个小清单,基于你的这些预测,我觉得大家为了取得成功应该去做哪些事。我们最后还会再聊,这里先说几条。第一,开始越来越多地用 Codex 或 Claude Code 来做你手头的工作,尤其要用里面的浏览器之类的工具。第二,让你的——让 agent 能够使用你的产品。如果你真的是做 SaaS 工具的,就让 agent 也能轻松地成为一个用户。第三,开始琢磨一个你能配合使用的 Slack bot,去试一试这些工具。我知道 Slack 自己也有个 Slack bot,听说也挺好用的,我还没玩过,但大家都很喜欢。所以去找一个能成为你公司内部 AI agent 的工具吧。还有,赶紧买 SaaS 股票。


[49:27]

[laughter]

(笑)


[49:28] Dan

Not investment advice. I think that's totally right. I will like my slight tweak is when you're thinking about building your software for agents, the current model is I'm building a CLI that an agent uses, but they're using it in a sort of like they're being I delegated a task to the agent and the agent's using the CLI. And what we what where I think it's going is you and the agent are using the app together. The agent's probably using the CLI, but you're using the web interface and they're they both need to be in sync. And that is I think a new challenge that's really interesting. Awesome. Anything else before we get to our next uh category? Bisas. That's the title.

这不构成投资建议。我觉得这些都完全正确。我稍微补充一点:当你在思考为 agent 构建软件时,现在的模式是——我做一个 CLI 给 agent 用,但 agent 用它的方式是「我把一个任务委派给 agent,agent 再去用这个 CLI」。而我觉得它会演变成的样子是:你和 agent 一起在用同一个 app。agent 大概还是用 CLI,但你用的是 web 界面,两边需要保持同步。我觉得这是一个全新的、非常有意思的挑战。太棒了。在我们进入下一个话题之前,还有什么要补充的吗?「Bisas」——这就是标题。


[50:14]

[laughter]

(笑)


[50:14] Lenny

Oh, man. Okay. So, the second uh category of predictions is around just the shape of the work that we're going to be doing is going to change. Uh what do you predict? There's all this interesting stuff in terms of in terms of the shape of work. Like once you're in this land where you've got, you know, these you've got async uh async agents off that you delegate work to then you've got your like Codex cloud code like work surface that that starts to happen. So, one thing that we see a lot internally and you also see this in the big model companies is the number of pull requests that you get is like skyrockets. You know, we have people, you know, in consulting or in ops roles or whatever who are or or editors just like making pull requests. Um and hey, that's really cool and it's a very different shape of work where you should you can expect that a higher percentage of your company or your users are going to be doing things that previously only technical users could do. And what that does is it creates all this pressure on the other end for the people who have to deal with all of the new code for how to deal with that. And so, I think there's a lot of there's a lot of interesting things that happen with that. Like so, for example, um uh like open claw, I mentioned that earlier. Pete gets like thousands of pull requests a day on open claw and then he has like and then he just spins up like 50,000 Codex instances and then sorts through them and then merges like a thousand of them. It's really crazy. I actually think that that's going to be more and more common. Um there's like it brings up a lot of really interesting questions around um which pull request should you merge? And you know, when you whenever you add capacity in one part of your process, like it breaks things. Um it used to be really hard to build things, and now it's very easy. So, the the point is not, can we build it? It's like, would it make sense with the rest of what we've built? And how do we keep a like sense of a coherent whole? And also, what do we delete? I think Entropic does this really well. Like they they delete a lot of stuff from Cloud Code to make sure that's not bloated. So, I I think there's a there's a lot of that going to happen. On one side, there's a lot of um non-technical people can do technical work, and then technical people are in charge of making sure that that work gets into a product or into a process in a cohesive, coherent way. And also, their product people are going to be doing that, too. And I think that's that's quite cool. Something I'm hearing from people is that now that everyone can do everything, like engineers can design, PMs can code, marketing people can ship stuff. There's just this like confusion about what the hell is my job anymore. Yeah. What am I responsible for, exactly? Like, am I supposed to be shipping stuff? Am I still a marketing person? And it's just creating a lot of confusion and certainty in the world. I think I think that's for real, and one of the things that I think is special about every is everyone is sort of a generalist and really loves like having their fingers in a lot of different pots, or whatever the metaphor is. I think that'll probably settle down at some point, and it'll feel more normal. Like, marketing people are still going to do marketing, even if they're touching the website. Like, that's just part of marketing now. But I also think that you can get a lot further being a generalist now, and that's like really cool, especially for for smaller companies. The The other thing that I think is interesting is there are definitely some new job roles that are a thing. And the thing that is becoming really clear is the whole forward deployed engineer concept I think is for real. And it comes out of every agent needs a human. Uh even like you go to the big model companies, they have they they have these agents that run internally. They have like teams of people that run these agents, you know? And I I don't think those teams are going away. The models are going to get more powerful, the agents are going to get more powerful, and the number of agents is going to grow, but people are still going to manage them. And so that looks like a very specific kind of person. And you know, we have a couple of those people internally here, and it's like the the people who are in charge of making sure your agents are working and doing the right thing. We also do consulting, so we we we lend that out to people, and and I think that's a big um that's a big thing that that people want, and it's another one of those places where you're like, "Hmm, automation was supposed to take away jobs, but it looks like it just created one or many."

好,第二大类预测是关于「我们将要做的工作,其形态本身会发生变化」。你的预测是什么?这里面有很多很有意思的东西。一旦你进入这种状态——手里有一堆 async agent,可以把任务委派出去,还有像 Codex、Claude Code 这样的工作界面,事情就开始发生变化了。

(Dan)我们内部经常看到一个现象,大模型公司里也一样:你收到的 pull request 数量会暴涨。我们这边有做咨询的、做运营的、当编辑的人,他们都在提 PR。这其实挺酷的,而且是一种全新的工作形态——你会发现公司里、或者你的用户里,有更高比例的人开始去做以前只有技术人员才能做的事。但这会在另一端制造巨大压力:那些得处理这些新代码的人,要想办法消化它。这里面会发生很多有意思的事。比如我前面提到过 open claw,Pete 每天能收到上千个 PR,然后他就起 5 万个 Codex 实例去筛,最后 merge 进去差不多一千个。真的很疯狂。我觉得这种模式会越来越普遍。它带出来一个很有意思的问题:到底该 merge 哪些 PR?而且,每当你在流程的某个环节加足马力,就会把别的环节搞崩。以前做东西很难,现在很容易了。所以问题不再是「我们能不能做出来」,而是「它跟我们已经做的东西放在一起合不合理」,以及「怎么保持整体的连贯统一」,还有「我们该删掉什么」。我觉得 Anthropic 这点做得特别好——他们会从 Claude Code 里删掉很多东西,确保它不臃肿。所以这类事会大量发生。一方面,大量非技术人员能做技术活了;另一方面,技术人员负责把这些活儿以连贯、统一的方式落进产品或流程里。而且产品同学也会做这事,我觉得这挺酷的。

我从大家那儿听到的一个说法是:现在人人都能干所有事——工程师能做设计,PM 能写代码,市场的人也能上线东西,于是大家就很困惑:「那我这份工作到底还算啥?我到底负责什么?我是该去上线东西,还是我还算个做市场的?」这在行业里造成了很多混乱和不确定。

(Dan)我觉得这是真实存在的。Every 有一点我觉得很特别,就是每个人某种程度上都是 generalist,特别喜欢什么都插一手。我觉得这种状况最终会慢慢稳定下来,会变得更「正常」。比如做市场的人还是会做市场,哪怕他们去碰网站——那现在也只是市场工作的一部分了。但我也觉得,如今当一个 generalist 能走得比以前远得多,这点很酷,对小公司尤其如此。

另一件我觉得有意思的事是:确实冒出了一些新的岗位角色。现在越来越清楚的一点是,forward deployed engineer(前线部署工程师)这个概念是真成立的。它的逻辑根源在于:每个 agent 都需要一个人。你去看那些大模型公司,他们内部跑着这些 agent,背后有一整个团队在运营这些 agent。我不认为这些团队会消失。模型会越来越强,agent 会越来越强,agent 的数量也会增长,但还是得有人去管它们。这对应的是一种非常具体的人。我们内部就有几个这样的人,他们负责确保你的 agent 在正常工作、做对的事。我们也做咨询,会把这种能力外借给客户。我觉得这是大家非常想要的东西,也是又一个让你感叹「自动化本来应该消灭工作,结果它好像反而创造了一个、甚至好多个工作」的地方。


[54:54]

[laughter]

(笑)


[54:54] Dan

You know, um and there's a specific type of engineer that really loves, you know, Nitesh, who's one of our uh who who fits this. He's an AI engineer and he he fits the sort of forward deployed um category, and he's on our team. He spends most of his time actually talking to one of our agents in Slack. We have an agent internally called Claudie, which runs our whole consulting practice. And and he spends a lot of time in Slack. Like there's there is code, and he is using Claud code and other things like that, but a lot of it is just talking to it and being like, "Why did you do this dumb thing? Like let's let's fix that, you know?" Um and so there's certain kinds of engineers that I think love that and love having their hands on the latest thing, and also love making this like being that's like in in the works in a work space and it looks a bit different than more traditional building more traditional software. And your sense there is we're not going to we're not near a place where these agents don't need a human. You said that so many times now that agents need a human and there's kind of like the setup part and then there's the maintaining it forever part. It feels like both are important. Is what I'm hearing like this is going to be a job for a long time. AI is not going to get smart enough to just automate it you fully automate for a while. Yes, I'm simultaneously extremely AI filled extremely and very bullish on humans and the role of humans in making sure that AI is working well. Interesting. Okay, so the two kind of buckets here that you're talking about one is like the way I think I hear what you described earlier is this the pace of shipping software and everything is just increasing which also means there's so much more work reviewing all this sloppy output. I was just talking to a data science friend and he was saying how his team is just as data science team is just their job used to be do analysis, answer questions, see if this experiment was a good was a was positive. Now it's just everyone's doing that and they're sharing the results and they're and they're like no this is not correct and most of their job is now reviewing bad data science work. Which is a problem and it means that and the same thing is happening with engineers and it means that you need more like you actually need data engineers for this and you need data scientists and it means that you haven't set up the appropriate systems or agents to help you with this. So like the way that it works inside of the big mono companies for example like at least one of them has literally a data science bot that every single person in the org can query that is hooked up to their data warehouse that knows who's who so that it knows at the warehouse level like who has permission to access what. And so all of the basic questions because they're there's a team that sets up this bot. All of the basic questions that people might want to ask that it sometimes gets that might get wrong that they're constantly making sure it's getting it right. And so the data science team doesn't have to answer all of the like [ __ ] questions because there's another team building an agent that that that is set up to do that really well. But if the team didn't exist the data scientist would hate their lives. Yeah. It does though make the job maybe less fun cuz you're just sitting there you know gardening people's sloppy work versus

有一类工程师是真的特别喜欢这个。我们团队里有个叫 Nitesh 的人就属于这类,他是 AI engineer,正好符合 forward deployed 这一类。他大部分时间其实是在 Slack 里跟我们的某个 agent 对话。我们内部有个 agent 叫 Claudie,整个咨询业务都是它在跑。他花大量时间泡在 Slack 里。当然也有写代码、也在用 Claude Code 之类的工具,但很多时候就是在跟它对话,问它「你刚干的这蠢事是为啥?咱把它修一下」。所以我觉得有些工程师就特别爱这个,爱上手最新的东西,也喜欢这种「身处一个工作空间里」的感觉,它跟传统的、更偏搭建软件的工作看起来挺不一样的。

(Lenny)你的判断是:我们离「这些 agent 不再需要人」的那一天还很远?你已经说过好多遍 agent 需要人了。这里好像分两块:一块是搭建的部分,一块是长期维护的部分,感觉两块都很重要。我听到的意思是,这会是个长期存在的工作,AI 短期内不会聪明到能把它完全自动化掉?

(Dan)是的。我同时既是极度的 AI 信徒,又非常看好人、看好人在「确保 AI 正常运转」这件事上的角色。

(Lenny)有意思。好,你说的这两个大类,一个是——我理解你前面描述的,是软件上线的节奏整体在加快,这也意味着要花更多精力去 review 这些粗糙的产出。我刚跟一个做数据科学的朋友聊,他说他们整个数据科学团队,以前的活儿是做分析、回答问题、看某个实验是不是正向的;现在变成人人都在做这些、然后把结果发出来,他们就得说「不对,这个结果有问题」,现在他们的工作大部分变成了 review 别人做得很糟的数据科学工作。

(Dan)这是个问题。同样的事也发生在工程师身上。这意味着你其实需要更多——你需要真正的 data engineer,也需要 data scientist,还意味着你没有搭好相应的系统或 agent 来帮你应对这件事。比如在大模型公司内部是怎么运作的:至少其中一家真的有个 data science bot,组织里每个人都能去查询,它接到了他们的数据仓库,知道谁是谁,所以在仓库层面它知道谁有权限访问什么。因为有一个团队专门搭这个 bot,所以那些大家可能会问的基础问题——它有时会答错的——这个团队一直在确保它答对。这样一来,数据科学团队就不必去回答所有那些破问题了,因为另有一个团队在搭一个专门把这件事做好的 agent。但如果这个团队不存在,数据科学家会恨死自己的生活。

(Lenny)不过这确实可能让工作变得没那么有意思了,因为你就是坐在那儿给别人粗糙的产出「除草」,而不是——


[58:04] Dan

what I think is like it it can actually make the job better because for the data scientist you're now not dealing with all the silly requests. You're dealing with the deep the deeper questions that are harder for the the team who's dealing with all the basic requests and building an agent to do that. It's it's like filtering all that stuff out so you can focus. Here's a question I've been thinking about. I was not planning to talk about this but it's something that I've been thinking about. So the question is which product tech role is the least changed now. So like engineers 100% of code AI now. It's like a completely different job. Product management a lot of the you know PRDs are you don't have to write as much. You can ship code. You don't have to wait for people. Design the whole design process dead according to recent guests just like there's no time to do the whole design process very different role. Data science very different work now. There's marketing there's sales. So here's the question. What do you think is the least fundamentally changed role so far? Well one interesting thing is you know I don't know if this counts but like CEOs and investors it seems still very very optional whether or not they use this stuff. Mhm. It seems that way. I I I think the opposite is actually true. Like my experience we do a lot of this with senior executives and senior leadership teams. My experience is that your company's only going to go as far as your CEO goes in AI and it's not something you can delegate. You have to have your hands in it cuz you don't otherwise you don't have an intuition for it. But for a long time it has seemed like yeah, that's something that the people who are doing the work have to do but like I don't have to do that. Like I'll just tell them what to do. And And so I think if you're a CEO, you kind of can get away with your day looking very similar. I I think that will change rapidly at some point where it'll be like oh no, I'm like way behind but for now because or maybe even middle managers, like those kinds of people I think are are it's fairly similar. I think like maybe sales because it's so so impersonal.

我觉得它其实能让工作变得更好。因为对数据科学家来说,你现在不用去处理那些幼稚的请求了,你处理的是更深的问题——那些对「负责所有基础请求、并搭 agent 去应对它们的团队」来说更难的问题。它等于是把那些杂事都过滤掉了,让你能专注。

这里有个我一直在琢磨的问题,本来没打算聊,但我确实想了很久。问题是:现在哪个产品技术类岗位变化最小?比如工程师,现在代码 100% 都有 AI,完全是另一份工作了;产品管理,很多 PRD 你不用写那么多了,你能直接上线代码,不用等别人;设计——按最近几位嘉宾的说法,整套设计流程基本死了,根本没时间走完整流程,角色变得很不一样;数据科学,工作方式也大变;还有市场、销售。所以问题是:你觉得到目前为止,哪个角色从根本上变化最小?

(Dan)有个有意思的点是——不知道这算不算——CEO 和投资人,用不用这些东西似乎仍然非常「可选」。

(Lenny)看上去是这样。

(Dan)我其实觉得恰恰相反。我们跟很多高管、高级领导团队做过这类事,我的经验是:你公司在 AI 上能走多远,取决于你的 CEO 在 AI 上走多远,这不是能委派出去的事。你必须亲自下场,否则你对它根本没有直觉。但长期以来大家的感觉是「那是干活的人该做的事,我不用,我告诉他们怎么做就行」。所以如果你是 CEO,你大概能蒙混过去,让你的每一天看起来跟以前差不多。我觉得这早晚会迅速改变,到时候会变成「糟了,我已经落后一大截了」。不过现在嘛——或者也可能是中层管理者,这类人我觉得情况也相当类似。我猜可能是销售,因为它太「非个人化」了。


[1:00:11] Lenny

that's yeah, that's my vote. You know, there it's sort of creeping up in the kind of BDR like we can deal with a lot of, you know, BDR type of type queries. You're only talking to like people who actually want it. And you can do it for sales it's like it's so useful to to like do research. Like my favorite codex, like one of my favorite codex experiences is we're hiring a head of L&D. And I you know, we always put out a job post whatever but I was like I feel like there's this company called General Assembly in New York and they do like they've done really good technology education for a long time. And so I was like I feel like someone who is into who who who worked at General Assembly and is now into AI would be really good and I just like literally typed it into codex and then like went off and was doing something else and I came back and it found like this the perfect guy. It was like worked at General Assembly, was an instructor like is super AI pilled and follows me on Twitter. So I just DM'd him and then I had dinner with him. And it's like that's crazy. You know, that would have taken so long before. And super valuable for sales, for recruiting, all that kind of stuff. Yeah, sales is where my mind went. Like the top of funnel AI is helping a lot with sourcing and qualifying things like that. It feels like the the work of a salesperson is not fundamentally different. Yeah.

对,那也是我的选择。它在 BDR(业务拓展代表)这块确实在慢慢渗透——很多 BDR 类型的问询都可以交给 AI 处理,你最后只跟那些真有意向的人聊。

(Dan)对销售来说,它做调研特别有用。我最喜欢的 Codex 体验之一是:我们当时要招一个 head of L&D(学习与发展负责人)。我们当然照例发了招聘帖,但我心想,纽约有家公司叫 General Assembly,他们做技术教育做了很久,做得真的很好。我就想,找一个在 General Assembly 待过、现在又入了 AI 坑的人会特别合适。我就直接把这话敲进 Codex,然后我就去忙别的了,回来一看,它找到了一个完美人选——在 General Assembly 当过讲师、深度 AI 信徒,还在 Twitter 上 follow 我。我直接给他发了 DM,然后约了一起吃晚饭。这太疯狂了,搁以前这得花多久。对销售、对招聘这些事都超有价值。

(Lenny)对,我脑子里想到的也是销售。漏斗顶端这块,AI 在 sourcing、资格筛选这些事上帮了很大忙。但感觉销售这个人本身的工作并没有发生根本性变化。

(Dan)对。


[1:01:38] Dan

And customer support has fundamentally changed. So that's interesting sales. So far so good for the for those folks. Yeah. Okay. So maybe just summarizing some of the predictions in this bucket of just like the shape of the work, how it's going to change. What I'm hearing so far is it's going to be a lot more reviewing of other people's output as a part of the work. And then two, there's going to be a lot of like almost babysitting of AI agents to make them do the thing you want them to do for deploying and then just gardening them along the way, make sure they continue to do their work. Um anything else before we get into our third bucket? I would sort of split it into less babysitting agents and more your forward deployed team is trying to build a whole system that makes it so that people who have less knowledge can use that system without like doing something dumb. And that's like a really interesting engineering challenge. I think babysitting kind of makes it feel like it's yeah, you're just kind of like, you know, waiting for it to [ __ ] up and then fixing it or whatever. And you can you can that can be the case, but I think a lot of it is just extremely interesting engineering challenge of building a system for to enable everybody else in the organization to do what used to be a technical job. And then if you're not one of those people, like you're the data scientist or whatever, you can go a lot deeper with AI into like really important questions that eventually probably filter into the work that the, you know, the forward deployed engineering team is doing, but is like more generative and more new and and and you're you're dealing with harder questions. One other one last thing that I think is really interesting is I think that we will be reading way more AI generated writing in documents and emails and we will like it. And I think we're we will already we are already doing this in coding where we read plan documents. Like I don't want an engineer to handwrite a plan document. That would be very silly. It would be It would be obviously silly. Um and I think the same is true, you know, when we did our our uh quarterly planning for every at the end of 2025 we did it all with Notion agents. And we just had a bunch of Notion agents and or we had really one Notion agent and then we had a top-level company strategy and then we had everybody in the company just um talked to an agent and it asked them about what happened last year, how did it go, what were your goals, what what do you want to do this year, what are your metrics, it pushed back, and then it was like, how does it How does this relate to the overall company idea? Like all that kind of stuff. And then I got I got these like incredibly good AI generated like strategy reports or or plan like quarterly plans or for each part of each team. And then I could go in and be like, okay, who needs to Who's like Who needs to talk to each other? Like which teams need to talk to each other that like don't know they need to talk to each other? Um and uh you know, who's Which one of these is like like actually low quality or which one of these is high quality? Like all that kind of stuff makes it it makes it a lot easier to process. Um and I see that all the time now. Like I I I consistently get AI generated stuff and there is a difference between an AI generated document that's slop and not. And the slop one is it took them less time to make it than it takes me to read it. And they don't stand behind every line. So my expectation is, if you send me an AI generated document, I think that's great. And if we talk about it and it's clear you have no idea what's in it, like big no-no. Not allowed to do that. Um and I I think we this this aversion to AI-generated stuff that will go away because the kind of strategy document that GPT-5.5 can write when it's directed well by someone on my team is way better than like them just like dinking and dunking like like their fingers on the keyboard. Right. Like most people are really bad at writing strategy documents. So, the bar is low. Yeah. And and same thing with email. Like I most of my email is written by GPT-5.5 and Codex right now. And I would I honestly would prefer it to say that it's coming from GPT-5.5 and I may change it to do that. But I had this I had this experience the other day where I had this I had to send an email to um to one of our investors and I asked Codex like go do it and you like Codex knows to ask me and it usually does, but this time it didn't. And it just sent the email. And I didn't look at it at all. And I was like, [ __ ] And so, I went to my sent and looked at it and I was like, oh, this is exactly what I would have sent. And so, it's like it's pretty close to to that a lot of the time. Um it can be like a little over formal and there's a couple things that that it's just when you really think about it, most of your email is kind of it's not it's kind of wrote. It's kind of prosaic. It's kind of I I definitely want to be the one to think about what it should say, like what what it should say, but the actual sentences don't matter that much to me usually. Sometimes they do a lot. And this is coming from a writer. Like I care a ton about writing. I think that human writing is incredibly important. And I expect we only publish human writing. Well, actually, we publish a mix of human and AI writing, but we always label it. Um sometimes it's nice to have an AI co-author on certain things. Um I absolutely think that uh human writing is important and I think that the the the reaction or the aversion to AI writing is silly. It's such an interesting lens on that because when people think about AI writing, I think about social media and videos. And your point is internally, if you're just like working on planning and documents and email and things like that, like that is much less scary that it's AI written. And to your to your point, people are already doing this. You almost prefer it a lot of times cuz people are really bad Totally. We have this too for external stuff. Like we publish all these guides and the guides are often agent they're agent assisted and the agent is a co-author and they're intended to be read both by humans and by agents. And that's because like, if you're writing a huge informational thing, I mean, you do this all the time. Um, in order to like really apply it, the best way to do that is just like have your agent ingest it and remember the next time I'm, you know, doing pricing to like remind me of this guide and we'll go through it together or whatever. It allows you to operationalize uh, the ideas much better and it allows you to go much deeper because agents can read like 10,000 pages in like a second. And so you you can you talk to the human about the story and the stuff that matters and the core ideas and the agent has all the details that it can then apply for you when you need it. Awesome. Anything else in this category before we get into our final category? No. Okay, let's do it. So, the final bucket is just who will be successful in this AI future that we are approaching {slash} what should people be working on to be successful in this next year or two? I am super super bullish on PMs. And I know that your audience will probably love that. Um, but my my anecdotal case that has convinced me of this is we have this guy internally his name is Marcus and he runs Spiral, which is our writing app. Marcus is a PM by training. He He previously ran Axios Axios's writing product and was it was a PM and had a big team and it got to, you know, tens of millions in of revenue and ARR. And he took a year off that job and just got super AI pilled. And just learned how to use cursor basically really well. Now, I think he uses cloud code, but he was extremely cursor pilled for a long time. And he's I would call him like lightly technical. Um like knows what a database migration is. Like if he has to look at the code, I think he can understand it, but he's like I we never could have hired him to do this job even a year ago. But the coding models have gotten good enough that he can pair the kind of the technical knowledge he does have with his really spiky product sense and sense for writing and sense for users. And it's like it's so dangerous. Like he ships faster than almost anyone on the team and he has such a eye for every single user, every single conversation, like what does it mean and how do we collect it into a story about like where we want to go next and what are the issues we need to fix and like all that kind of stuff. And I think that he feels liberated cuz he doesn't have to organize a whole team of people to do that. He can just do it. And it's super impressive and it makes me very, very bullish on any PM who gets like really AI pilled. Music to my ears, Dan. You're making a lot of very happy listeners here. I've been saying this for a long time, too. It's just like the skills you need to build are the things like the building out is done for you. What do you need to be good at? Figuring out what to build, figuring out if it's great, figuring out what problems to solve. So, I love that you're actually seeing this come to fruition. I I I really believe it.

而客服则发生了根本性的变化。所以这点挺有意思的——销售目前还算安然无恙。

(Lenny)好。那我们稍微总结一下「工作形态如何变化」这一类的预测。我目前听到的是:第一,工作中会有大量去 review 别人产出的部分;第二,会有很多近乎「带娃」式地照看 AI agent,让它们去做你想做的事、把它们部署上线,然后一路「除草」、确保它们持续干活。在进入第三类之前,还有别的吗?

(Dan)我会换个说法:少强调「照看 agent」,多强调「你的 forward deployed 团队是在搭一整套系统,让知识较少的人也能用这套系统、又不至于干出蠢事」。这是个很有意思的工程挑战。我觉得「带娃」这个词会让人感觉你就是干等着它搞砸、然后去修——确实有这种情况,但很多时候其实是一个极其有意思的工程挑战:搭一套系统,让组织里其他所有人都能去做以前属于技术岗的活。而如果你不是那种人——你是数据科学家之类的——你就可以借助 AI 往真正重要的问题里钻得更深,这些问题最终大概率会反过来流进 forward deployed 工程团队的工作里,但它更具创造性、更新颖,你处理的是更难的问题。

还有最后一件我觉得特别有意思的事:我认为我们以后会读到多得多的、AI 生成的文档和邮件,而且我们会喜欢它。在写代码这块我们其实已经在这么做了——我们会读 plan 文档。我可不想让工程师手写一份 plan 文档,那太蠢了,明显很蠢。我觉得别的也一样。2025 年底我们给 Every 做季度规划时,全程都用 Notion agent。我们就用了一堆 Notion agent——其实主要是一个 Notion agent——然后定了一个公司层面的顶层战略,接着让公司里每个人去跟 agent 对话,agent 会问他们:去年发生了什么、进展如何、你的目标是什么、今年想做什么、你的指标是什么,还会反驳你,然后问「这跟公司整体的想法怎么对得上」之类的。最后我拿到的是一份份质量好得惊人的、AI 生成的战略报告或季度计划,每个团队、每个部分都有。然后我就能进去看:谁跟谁需要对齐?哪些团队之间需要沟通、但他们自己还没意识到需要沟通?这里面哪个其实质量很差、哪个质量很高?这一切都让处理起来容易太多了。

我现在到处都能看到这种东西。我经常收到 AI 生成的内容,而 AI 生成的文档是「水货」还是不是「水货」,是有区别的。所谓水货,就是他们生成它所花的时间,比我读它所花的时间还短,而且他们对里面的每一行字并不负责。所以我的期望是:你给我发一份 AI 生成的文档,我觉得很好;但如果我们一聊起来,明显你根本不知道里面写了啥,那就是大忌,绝对不允许。

我觉得这种对「AI 生成内容」的反感会消失的,因为当我团队里的人把 GPT-5.5 引导好之后,它写出来的那种战略文档,比他们自己噼里啪啦在键盘上敲出来的要好得多。对吧——大多数人其实都写不好战略文档,所以门槛本来就很低。

(Lenny)对。

(Dan)邮件也一样。我现在大部分邮件都是 GPT-5.5 和 Codex 写的。老实说我甚至更希望它注明这是 GPT-5.5 写的,我可能会改成这样。前几天我有过一次经历:我得给我们一位投资人发封邮件,我让 Codex 去搞定。Codex 一般会知道要先问我、通常也会问,但这次它没问,直接就把邮件发出去了,我压根没看。我当时就想「妈呀」。然后我去「已发送」里翻出来一看——咦,这正是我自己会发的内容。所以很多时候它已经相当接近那个水平了。它有时会稍微过于正式、还有几个小地方——但你真去想,你大部分邮件其实都挺套路、挺平淡、挺程式化的。我绝对希望由我来想「它该说什么」,但具体那些句子怎么遣词,对我来说通常没那么重要;当然有时候非常重要。

而且这话是从一个写作者嘴里说出来的——我极其在意写作,我认为人类的写作无比重要。我们只发表人类写的东西……好吧,其实我们发表的是人类写作和 AI 写作的混合,但我们一定会标注清楚。有些内容有个 AI 合著者会很不错。我绝对认为人类写作很重要,但我觉得那种对 AI 写作的反应、那种反感是很傻的。


[1:10:50] Lenny

This could be the highest rated podcast episode of my whole podcast. There is going to be like

(接前面的话题)这或许会成为我整个播客里评分最高的一期。这下要……


[1:10:53]

[laughter]

(笑)


[1:10:53] Lenny

Hell yeah. It's going to be okay. Stats is Stas is back, PMs are back, you know. This is the most contrarian episode I've ever done.

太棒了。会没事的。Stats——Stas 回来了,PM 也回来了,你懂的。这是我做过的最反共识的一期节目了。


[1:11:03]

[laughter]

(笑)


[1:11:04] Dan

Oh my god. So, okay, so the other the other people that I think are going to be like super super power people and I again I this is cuz we see this internally is full stack designers. If you're a designer and you're in these tools all the time, you're so used to um okay, I make this beautiful interaction and the engineer like just doesn't want to do it or it doesn't like happen the way I think it should happen or you know, there's all this stuff and I see so many designers for us internally or externally where they now feel so empowered to like go build stuff cuz they're like I have all these ideas to make things look amazing and these interesting interactions and that's the exact thing that it's really hard to do with live coding because it just all looks the same so it all looks like slop and they can make stuff that looks so different and now they can actually build it. And what you see when we work with them internally is now they're just like they're just making pull pull requests. Like they don't they don't need to hand it off as much. Sometimes they do but like a lot of times they just make pull requests and it's like the thing is built and that's it and I think that's incredible for the way that companies work but it's also there's a huge opportunity for those people to become much better and like start their own thing cuz they can they can make stuff now and I think designers are such creative people and I think AI is like a super tool for anyone like that. I so agree. Even though there is cloud design, there's all these AI designing tools, like once you see it, you're like that's definitely cloud design. And they're like the creativity to your point is it just feels like it's going to be more and more valuable to do to stand out from all the slop that people are shipping and launching all constantly. So, I completely agree. It's It's interesting that designer roles I do I do research on the job market and interestingly designer roles have not grown in a while. So, I'm waiting to see if that becomes a big trend just like we need more designers. Hm, that is really interesting. We'll see.

天哪。好,那另一类我觉得会成为超级超级强者的人——我之所以这么说,是因为我们内部就亲眼见到了——就是全栈设计师。如果你是个设计师,整天泡在这些工具里,你太熟悉那种感觉了:好,我做了一个特别漂亮的交互,结果工程师根本不想实现,或者实现出来跟我设想的不一样,诸如此类一堆破事。我看到我们内部还有外部好多设计师,现在都感觉特别被赋能,能自己动手去搭东西了,因为他们心里装着一堆点子,想把东西做得超好看、做出各种有意思的交互。而这恰恰是 vibe coding 很难做到的,因为它做出来全都长一个样,全是 slop。但设计师能做出风格迥异的东西,而且现在他们真能把它给搭出来。我们内部跟他们合作时看到的就是,他们现在就是直接提 PR,不太需要再交接给别人了。有时候还是要交接,但很多时候他们直接提个 PR,东西就建好了,就这么简单。我觉得这对公司的运作方式来说太了不起了。同时对这些人来说也是个巨大的机会,他们可以变得厉害得多,甚至自己出来创业,因为他们现在真能做出东西了。我觉得设计师都是特别有创造力的人,而 AI 对这类人来说简直就是个超级工具。 >> 我太同意了。虽然现在有 AI 设计、有那么多 AI 设计工具,但你一眼就能看出来,那肯定是 AI 设计的。所以正如你说的,创造力——为了从大家不停产出、不停上线的那堆 slop 里脱颖而出——感觉只会越来越值钱。所以我完全同意。有意思的是,我会研究就业市场,而有趣的是设计师岗位已经好一阵子没增长了。所以我在等着看,这会不会变成一个大趋势,就是大家突然都说我们需要更多设计师。 >> 嗯,这确实挺有意思的。咱们走着瞧。


[1:12:57] Dan

Yeah. We'll see. We'll see. That that might be a way to predict this is are people hiring more designers? I don't know. That is interesting. Yeah. All right. Uh so, that's so PM designer thriving. PM designer thriving. Um I also just think generally the AI job apocalypse is not really a thing. Absolutely, we see companies starting to reorganize and I think that makes a lot of sense. I I think to be honest a lot a lot of the reorganization you can say it's AI, but it's like we over hired and like the company's not doing as well and all that kind of it was like coming and this is a good excuse. But the like mass unemployment thing I think that like some AI CEOs are talking about like I think that's not going to happen. The the pattern that I see so far and again, I don't have a total crystal ball, but I I do feel like we've seen enough of the new model drops to like have some sense of how this is going is that what a new model drop does or what models do in general is they make yesterday's human competence cheap. So, what I mean by that is they ingest all this data of what what has happened already and they make it really cheap to deploy that in in whatever situation you want as your as your own, right? Um and what happens then is every this is a new this is a new power that everyone has. So, it gets adopted super rapidly and it's and suddenly that stuff is everywhere. It's like suddenly anyone can make a landing page, there's new landing pages everywhere. Suddenly everyone can write, there's like slop tweets everywhere. But what's interesting is because it's all from because it's all coming from these models and everyone's using basically the same models, uh it all looks the same if you use it in the in the most default basic way. And so, that's it becomes commoditized. Like it's not valuable anymore. And what humans do is we sort of go in there and we're like, "Yeah, we have all this like frozen human competence from yesterday. How do I use this like make something new and interesting?" And I really think that structurally, because of the way the models work, because of the financial incentives of model of model companies to like make them um uh compliant and aligned, structurally, there are always going to be trailing behind those people who are taking taking the models and using them to make new expertise or or make new things that haven't been done that way before for their very very particular situation. And that stuff is going to get incorporated into the models, but again, it will create room for people to um to push further ahead. And I think that you see this in a small way in like pretty much all the jobs is like engineers. Suddenly, everyone's an engineer. That doesn't mean we fire the engineers. There's like way more demand for engineers cuz you need the engineers to like figure out, "Okay, this is all slop. How does this actually How should this actually go in our code base?" And I think that's something that the benchmarks rising don't doesn't really capture. And uh it feels like a thing that will take a long time to change. People may be hearing in this uh prediction here of just, "Okay, the job apocalypse is not going to People are not going to be all fired. There's going to be human jobs remaining for quite a while." It may be almost too comforting because you may you probably have to change the way you operate to still have a job in the future. Do you have any sense of just like, "Here's what you need to do to not be one of these layoffs?" Yes. And I think that is actually super important. Um the only thing you need to do is ride the models. And that means use them for whatever it is that you do. You know, we've talked about how Codex and Co-work are becoming the a of standard operating system for work. If you're just doing that and when new models come out, you're trying them and figuring out, okay, how can I now they're new powers, how can I use them instead of just being like, I'm going to like try to ignore it cuz it like makes me afraid, which I think is honestly it's rational, it's a reasonable response. And also uh if you ride on top of them, they ex- extend your powers in a way that doesn't leave you behind. Like you you're you're you're part of the future and part of the way work happens and I think that uh we're going to need people doing that for a very, very long time. I like this term ride the model. So, the what's like saying you all comes out, what do you think someone say working at I don't know, Salesforce. Say a PM at Salesforce, what should they do to ride the model? Well, one of the things that's really interesting is a lot of companies like handicap their employees from even doing this because like I don't know what model I don't know if you can use the latest models at Salesforce, you know, like a lot of times you have to wait or it's, you know, whatever. So, maybe you have to do it on your in your off time. But, the thing that I really like to do with new models is play. And there there are there are certain things where I know it can't quite do it yet, but when a new model comes out, I like always turn the rock over again to be like, can it do it now? You know, um so, it you know, it could not do the senior engineer benchmark last time and I turned it over turned the rock over again and now it's out of 60 out of 100, which is like really good. Um so, the way to ride the models is like not one specific thing cuz they're always changing, but it is to be curious and playful, to apply the model the new model to whatever it is that you care about, whether that's your job or something outside of your job and to keep turning over rocks uh because it may not work now, but it may work eventually, it probably will work eventually and the way that you use it matters. So, what's really cool is that I think people think of the edge of AI as being in San Francisco. And I actually don't think that that's where it is. I think the edge of AI is wherever AI meets like a real human doing something. Because the people in San Francisco, they're making it, but they don't actually know a lot about how to use it. They don't know or at least they don't know everything about how to use it. They need to see how other people use it. And so you whenever new model comes out, you get to be one of the first person one of the first people in the world to discover what it might be useful for. And like that's it's like a new discovery. And I think that's why, for example, we're in we're in Brooklyn. But I I really think of us and I think we are like quite far ahead of people in San Francisco because we just use them for everything. And um if people uh if people do that consistently, I think it's going to be very hard to lose. That is one of the amaze most amazing things about AI right now is no matter how much money you have or little money you have, you have access to the most advanced AI model. Like it's not free, so you need some money. Uh uh but like and you can get it immediately when it comes out. Maybe the only people that have an advantage are the people working at OpenAI or Anthropic. Um but otherwise, it's just like available. I know I was at I was at uh their event with you their code event with you last week and um or a couple weeks ago and they're they're like all using mythos and I'm like, "God damn it." So annoying. [laughter] But I I think that's totally true. Like that is if IBM had invented AI, you can bet it would not be like this. And it would be like a bajillion dollars and only like the top companies could use it and they would be using it in the in the weirdest, most uninteresting ways. And I think there's it's there it's really important that AI was built in America and in the Silicon Valley culture that's like we want to make intelligence too cheap to meter. Like that's not the default stance. And um it means that everyone has this broadly accessible tool that they can use and I think that's amazing. It's such a good point and interestingly it's also created the most fastest growing companies in history, the biggest companies in history. That's true. Not the way to

对,走着瞧。走着瞧。这或许能用来预测这件事——就是看大家有没有在招更多设计师。我也不知道。但确实挺有意思。好。所以 PM 和设计师都在蓬勃发展。PM 和设计师蓬勃发展。我还觉得,总体上看,所谓 AI 引发的就业末日其实根本不存在。我们当然看到一些公司开始重组,我觉得这很合理。说实话,很多重组你可以说是因为 AI,但其实更像是当初招人招多了、公司业绩又不太行,本来就该裁了,而 AI 正好是个好借口。但某些 AI 公司 CEO 嘴里那种大规模失业的说法,我觉得不会发生。我目前看到的模式——我当然没有水晶球——但我确实觉得我们已经见过足够多次新模型发布,对事情会怎么发展有了点感觉:新模型发布、或者说模型整体上做的事,就是把昨天的人类能力变得很便宜。我的意思是,它们把已经发生过的所有数据吃进去,然后让你能极其廉价地把那种能力部署到任何你想要的场景里,当成你自己的能力,对吧。接着发生的是,这成了每个人都拥有的新能力,所以它被超快地采纳,一下子到处都是。突然之间人人都能做落地页,于是新落地页满天飞;突然之间人人都能写作,于是 slop 推文满天飞。但有意思的是,正因为这些全都出自这些模型、而大家用的基本是同一批模型,所以你要是用最默认、最基础的方式来用,做出来的东西全都长一个样。于是它就被商品化了,不再值钱了。而人类会做的,是钻进去说:好,我们手里有这一整套昨天冻结下来的人类能力,那我怎么用它去做出点新的、有意思的东西?我真心觉得,从结构上讲,由于模型的工作方式、由于模型公司在财务上有动力把模型做得顺从、对齐,它们在结构上永远会落后于那些拿模型去创造新专长、或针对自己非常非常具体的情况做出前所未有之事的人。那些东西会被吸收进模型里,但同样,这又会给人留出空间去往前推进。我觉得几乎所有工作里都能看到这种小规模的体现,比如工程师。突然人人都是工程师了,但这不意味着我们就把工程师裁了。对工程师的需求反而大得多,因为你需要工程师来搞清楚:好,这些全是 slop,那这玩意儿到底该怎么进我们的代码库?我觉得这正是 benchmark 上涨所没法真正捕捉到的东西,而且感觉这是个要花很长时间才会改变的事。 >> 大家听到你这个预测——就是好,就业末日不会到来、不是所有人都会被裁、人类的岗位还会存在相当长一段时间——可能会觉得这几乎太让人安心了,因为你将来要想保住工作,很可能还是得改变自己的做事方式。你有没有什么具体的感觉,比如:要不被裁掉,你得做这几件事? >> 有。我觉得这其实特别重要。你唯一要做的,就是骑在模型上往前走(ride the models)。也就是说,不管你干的是什么,都拿它们来用。我们聊过 Codex 和 Co-work 正在成为工作的标准操作系统。如果你就在这么干,并且每当新模型出来时你都去试、去琢磨——好,它们现在有了新能力,那我现在能怎么用——而不是因为它让你害怕就想着干脆无视它。说实话,那种害怕是理性的、是合理的反应。但你要是骑在它们上面,它们就会以一种不把你甩在身后的方式拓展你的能力。你就成了未来的一部分,成了工作运转方式的一部分。我觉得这种人,我们会需要非常非常长的时间。 >> 我喜欢 ride the model 这个说法。那比方说,假设新模型出来了,你觉得一个在——我不知道——Salesforce 工作的人该怎么办?比如 Salesforce 的一个 PM,他该怎么骑在模型上? >> 有意思的一点是,很多公司其实把员工的手脚都捆住了,连这事都没法干。因为我不知道你在 Salesforce 能不能用上最新的模型,对吧,很多时候你得等,或者反正就是各种限制。所以也许你得在自己的业余时间干。但我用新模型最喜欢做的事就是玩。有些事我知道它现在还做不到,但每当新模型出来,我总会再把那块石头翻过来看看:现在能做到了吗?比如上次它过不了那个资深工程师的 benchmark,我又把石头翻过来,现在它能拿 100 分里的 60 分了,已经相当不错了。所以骑模型这事没有某一个具体的招,因为它们一直在变;它的关键在于保持好奇和爱玩,把新模型用到你在意的任何事情上,不管是工作还是工作之外的东西,然后不停地翻石头。因为它现在也许不行,但最终也许就行了,而且很可能最终就是行的,而你怎么用它,是有讲究的。所以特别酷的一点是,我觉得大家都以为 AI 的前沿在旧金山。但我其实不觉得前沿在那儿。我觉得 AI 的前沿在任何 AI 遇上一个真实的人去做真实事情的地方。因为旧金山那帮人,他们在造它,但他们其实并不太懂怎么用,或者至少不是什么都懂。他们需要看别人怎么用。所以每当新模型出来,你都有机会成为全世界最早一批发现它能派上什么用场的人。这就像一次全新的发现。我觉得这也是为什么——比如我们在布鲁克林——我真心觉得我们其实远远走在旧金山那帮人前面,因为我们什么都拿它来用。如果大家能持之以恒地这么做,我觉得就很难输。 >> 这正是 AI 现在最了不起的地方之一:不管你钱多钱少,你都能用上最先进的 AI 模型。当然它不是免费的,所以你得有点钱。但你能在它一发布就立刻用上。也许唯一有优势的人就是在 OpenAI 或 Anthropic 工作的人。但除此之外,它就这么随手可得。我上周——其实是几周前——跟你一起参加了他们的活动,他们的那个代码活动,他们都在用 Mythos,我当时就想,靠。太气人了。(笑)但我觉得这完全没错。如果当年是 IBM 发明了 AI,你敢打赌它绝不会是现在这样。它会贵得离谱,只有顶级公司用得起,而且他们会用最古怪、最无趣的方式去用。我觉得 AI 诞生在美国、诞生在硅谷那种「我们想让智能便宜到不值得计量」的文化里,这一点真的非常重要。这并不是默认的姿态。它意味着每个人都有这么一个广泛可及的工具可以拿来用,我觉得这太了不起了。 >> 说得太好了。而且有意思的是,它还顺带造就了史上增长最快的公司、史上最大的公司。 >> 没错。


[1:20:59] Lenny

Those Silicon Valley guys, they're they're smart. If I zoom out on the conversation, it's really interesting. There's a kind of these two sides to the coin. One is not a lot is actually like so much is not changing. SaaS continues, jobs not disappearing. We're still emailing each other. We're still working in Slack. Like a lot of the work not changed. On the other hand, every role transformed. Engineers don't write code. PMs don't write PRDs. Design and design, you know, it's like it's so interesting how much has changed, how much has not changed. I don't know. It's interesting that people think it's going to be this whole new world, but in many ways it's okay. It'll continue the way it is with a lot of stuff around the edges. That's that's how I feel. Like I'm simultaneously so excited and it feels like everything has changed. And I'm so bullish on it and and the and the progress that we're going to make and all that kind of stuff. And yeah, I just I feel like there are there are these things where they're going to be pretty similar to how they are and that's probably good. And I think generally our intuitions about the future the the model that I have of what our intuitions are about the future is the intuitions that people had in the Middle Ages about like what happened at the end of the horizon, you know, it's like are there dragons? Like does it drop off into nothingness or whatever? You know, like a lot of people have a lot of deep intuition that there's something terrible going to happen over the horizon. And also that uh some people are like there's something incredible. It's It's to change everything. We're going to all all going to be happy as a utopia. And what happens is you get there and you're like, there's some really cool things, there's some not cool things, and it's just another horizon. And I think that's that's the way to think about the future. And until you get to that place where you're starting to see it, and I think we get to see it cuz we get to see it internally all the time, it's important not to let your your mind get away from you and being like, this is going to happen and this is going to happen and whatever cuz you're you're going to tell a story that sounds sounds so real in the moment, but um later on you're like, actually it's much more complex than that and somewhere it's sort of a both everything has changed and nothing has. Um and once you get there, I think you're you're sort of start starting to see like, oh yeah, this is a real thing. Part of it is that the AI companies are very good at scaring us about what might might happen in the future. And I think that's actually shifting. I think that they've realized maybe we should not freak everybody out about the dangers.

那些硅谷的家伙,他们是真聪明。如果我把这场对话拉远来看,会发现特别有意思。这事儿有点像一枚硬币的两面。一面是,其实没那么多东西在变——好多东西根本没变。SaaS 还在,工作没消失,我们还在互相发邮件,还在 Slack 里干活,很多工作并没有改变。但另一面是,每个角色都被重塑了。工程师不写代码了,PM 不写 PRD 了,设计师也是……你看,变了这么多,又没变这么多,真的特别有意思。我也说不好。有意思的是大家都以为会迎来一个全新的世界,但很多方面其实还好,它会基本照旧运转,只是边边角角多了一堆新东西。 >> 我就是这种感觉。我同时既无比兴奋、觉得一切都变了,又对它无比看好——看好我们将要取得的进展,等等这一切。但同时,是的,我就是觉得有些东西会跟现在差不多,而那大概是好事。总体上我觉得我们对未来的直觉……我对「我们对未来的直觉」的模型是这样的:它就像中世纪的人对地平线尽头会发生什么的那种直觉,你懂的,那边有龙吗?是会掉进虚无吗,还是怎样?很多人都有一种很深的直觉,觉得地平线那头会发生什么可怕的事。也有些人觉得那边有什么不可思议的好事,要改变一切,我们都会幸福得像活在乌托邦里。而实际情况是,你真走到那儿一看,会发现有些挺酷的东西,有些不那么酷的东西,然后它就只是又一道地平线而已。我觉得这才是看待未来该有的方式。在你走到那个开始能亲眼看见的地方之前——我觉得我们能看见,因为我们内部一直在看见——很重要的一点是别让脑子脱缰,别去想这会发生、那会发生、blah blah,因为你会编出一个当下听起来无比真实的故事,可后来你会发现,其实比那复杂得多,而真相在某处是「一切都变了」和「什么都没变」同时成立。一旦你走到那儿,我觉得你才会开始看清:哦对,这是个真实的东西。其中一部分原因是,AI 公司特别擅长拿未来可能发生的事来吓唬我们。我觉得这一点其实正在转变。我觉得他们意识到,也许我们不该把所有人都吓得对那些危险惊慌失措。


[1:23:20] Dan

I that PR strategy just does not make any sense to me. I I do think that it's like genuine, but it's so ineffective and um and I I think it's also wrong. Mhm. How about we um end with maybe just like a few things listeners should do to be successful over the next year with the the way the world is moving. Write the models. I would uh try all of your workflows in Codex or Co-work and see how that works. And if your company doesn't let you do it on your own time, I would try out some of these um agent products like Open Claw or Hermes or um for less technical people there's there's like Victor, we have 1 + 1's. I I would get comfortable with both of those ways of working. And try to like try to have fun. I think there's too much of I'm doing this because I have FOMO like it might I might lose my job or like I might miss out on this big thing or whatever and the best way to actually figure out interesting useful things to do with AI is to like do something enjoyable. We had a um Nikhil Singhal was on the podcast and the way he described it is you got to find your moment of joy with AI. Once you find like, "Wow, I can't believe AI did this for me. This is awesome. I'm going to keep building stuff." Yeah, I agree.

那套 PR 策略我是真没法理解。我确实觉得它是真心的,但它太没效果了,而且我觉得它也是错的。 >> 嗯哼。 >> 那我们要不就这样收尾吧——就讲几件听众可以去做的事,让他们在接下来这一年里、在世界这么变的当下能做得不错。骑在模型上。我会建议你把你所有的工作流都拿到 Codex 或 Co-work 里去试一遍,看看效果怎样。如果你公司不让你在自己的时间里干,我会建议去试试这些 agent 类产品,比如 Open Claw 或者 Hermes;给不那么技术的人,还有像 Victor 这种,我们也有 1+1 之类的。这两种工作方式我都会建议你去熟悉。还有,尽量去找乐子。我觉得现在有太多人是出于 FOMO 在做这件事——怕丢工作、怕错过这个大风口什么的——而真正想搞清楚 AI 能干出哪些有意思又有用的事,最好的办法就是去做点让你享受的事。我们之前有 Nikhil Singhal 上过播客,他的说法是:你得找到你跟 AI 在一起的那个快乐时刻。一旦你找到那种「哇,我简直不敢相信 AI 帮我做出了这个,太牛了,我要继续接着搭东西」的感觉。 >> 嗯,我同意。


[1:24:43] Lenny

you haven't seen that yet, then it's just like try find try solving it. The thing I hear a lot is just find a problem in your life or work and see if AI can do it. Go to loveabull, go to clockcode, go to replit. Try to build the thing and often it's like, "Holy [ __ ] this is so cool." Dan, is there anything else that we haven't covered? We've gone deep on so much. Is there anything else you want to share? Anything else you want to predict or just say before we get to our very exciting lightning round? Uh I think we covered it. We we did a lot. This is This is awesome and I'm very excited to see how well or poorly I do uh in a year and I hope that you hold me to it. We're going to We're going to have AI score us. How about that? We'll Look look at the world like a dance prediction series. Well, with that Dan Shipper, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? I'm ready. What are two or three books that you find yourself recommending most to other people? Um obviously Annie Dillard. Um I Everyone at Every has to read The Writing Life. Like when you join you get a copy and you have to read it. Uh you only have to read the last chapter though. I think the last chapter is incredible and it is at the intersection of writing, technology, and the future and it's like it's relationship to the future and to time and I think that's like it's it's everything about every like wrapped up into like a very tight chapter. It's so good and I think Annie Dillard just generally is fantastic. What else do I recommend? I'll just I'll just tell you a couple things that I've read that I like really liked recently um and and whenever I like something I always just like tell everyone about it. So, um, I have recommended these a lot. Um, I I've been I've been reading One of the things I I learned, which I didn't know, is Churchill's a really good writer. And he has a whole history of World War II that he wrote, and it's like a combination history and memoir. And I think that's so cool because he was there, you know, he did it. And there's something about what we do at everywhere. I I feel some like sort of kinship with that of like we're building stuff, we're writing stuff, and it's very rare to find people that also do that. And and so, Churchill's history of World War II is fantastic. I just finished the first volume. I'm on the second volume. The Nazis just invaded France. Very It's very captivating stuff. Um, so, that's one. I also just I I've been on a like a little bit of like a quantum physics like kick recently. AI is very actually very good for quantum physics if you get into it. And there's this book called The Rigor of Angels that I just finished, which is um, it's like a it's a history of ideas that relates uh, Heisenberg, who has the his uncertainty principle, um, Borges, who's uh, uh, uh, like a uh, uh, Argentinian uh, fiction writer is wrote a bunch of great short stories that are actually starting to get like a lot of play now cuz they're very AI-related, and um, and Kant. And very cool, like super mind-blowing. Lots of like interesting overlaps with AI stuff. And uh, yeah, highly recommend. I feel like we could have a whole podcast episode about your reading and uh, books you recommend. I know this is a a passion of yours. My current obsession is The Power Broker. I don't I think we talked about it when I was visiting you. It's just never ends, but it's uh, surprisingly compelling to read through the history of New York. Okay, second question. Do you What is a recent movie or show you really see recently enjoyed if you have time for TV? So, I've been watching a lot of basketball, so that's one. Um I'm I became a Knicks fan like this this year, so uh that's really fun. But uh I recently watched this I guess it's like a it's like a mini-series documentary called The Dark Wizard about this guy Dean Potter who he was like Alex Honnold before Alex Honnold was Alex Honnold. And uh he just has this like very extreme personality where he's like free soloing everything and then he's like, you know, base jumping in in like a wingsuit and stuff like that and it's sort of exploring his psychology and what happened to him and um I I don't know. I I kind of like stuff like that. Like there's another one called 100 Foot Wave where it's like about people who are trying to like big wave surfers. There's something about that that's sort of I guess it just reminds me of founders or whatever, but um The Dark Wizard, highly recommend. Is there a product you recently discovered that you really love? Codex. It's like it's the best It's really good. It's really good. Do you have a favorite life motto that you often come back to in work or in life? Yes, I have several. Um the the like the core one that I wrote for myself in college was um do things worth writing about and write things worth reading. And uh and then there's there's this guy Rob Brezsny who's like very um very popular in like you know, the the AI meditation like overlap discourse, which is also a big thing. Um and who I also I really like him. He's dead, but I think he's amazing. And I listened to like so many of his talks and there's like this one talk that he gives where it's just like one sentence, but he just talks about like when you're dealing with stuff that's hard, what you want to do is be able to relate to it from a position of spaciousness and strength. And there is something I think really interesting and important in that. Like a lot of the meditation discourse or just generally like how do you deal with hard things? It's like a little bit more of like the David Goggins like you just got to like just got to like go for it kind of and like just um and sometimes that sometimes that can work. And also I think sometimes when you're dealing with things so for example when you're dealing with I'm super afraid of like how AI is going to um um you know, change my job. It is it has been very helpful for me to be like am I coming at this from a vantage point of spaciousness and strength and if not can I like get there? Because it will be much more productive for me to deal with it from that place. And that has been very very helpful for me. Wow. I love that. Well, our final question uh just on the on the theme of this conversation. Curious if there's just like an AI tool that you think is still kind of underrated that you're just like recently uh I mean

如果你还没体验过那种感觉,那就去试着解决一个问题、去找到它。我经常听到的一个建议就是:在你的生活或工作里找一个问题,看看 AI 能不能搞定它。去 Lovable、去 Claude Code、去 Replit,试着把那东西搭出来,往往结果就是「我靠,这也太酷了吧」。Dan,还有什么我们没聊到的吗?我们已经聊得很深了,还有什么你想分享的?在我们进入特别精彩的快问快答之前,还有什么想预测或想说的吗? >> 我觉得都聊到了。我们聊了好多。这太棒了,我特别期待一年后看看自己预测得有多准或多离谱,希望你到时候来验收我。 >> 我们会让 AI 来给我们打分,怎么样? >> 好。 >> 把世界当成一个 Dan 预测系列来看。那么,Dan Shipper,我们就此进入特别精彩的快问快答环节。我有五个问题给你,准备好了吗? >> 准备好了。 >> 有哪两三本书是你最常推荐给别人的? >> 显然是 Annie Dillard。Every 公司每个人都得读《The Writing Life》。你一入职就会拿到一本,必须读。不过你只需要读最后一章。我觉得最后一章太精彩了,它正好处在写作、技术和未来的交叉点上,讲它与未来、与时间的关系,我觉得它就像把 Every 的一切都浓缩进了一个非常紧凑的章节里。太好了,而且我觉得 Annie Dillard 整体上都棒极了。我还推荐什么呢?我就跟你说几本我最近读了特别喜欢的吧,而且我只要喜欢一样东西就总爱逢人就安利。这几本我推荐得很多了。我最近一直在读——我学到一件我之前不知道的事——丘吉尔其实是个很好的作家。他写了一整部二战史,是历史和回忆录的结合体。我觉得这太酷了,因为他人就在现场,是他亲身做的。我们在 Every 做的事,我对此有种说不清的亲切感:我们在搭东西、在写东西,而能同时做这两件事的人非常罕见。所以丘吉尔的二战史棒极了。我刚读完第一卷,正在读第二卷,纳粹刚入侵法国。非常引人入胜的东西。这是一本。另外我最近还有点迷上了量子物理。如果你钻进去,会发现 AI 其实特别适合学量子物理。有本书叫《The Rigor of Angels》,我刚读完,它是一部观念史,把几个人联系在一起:海森堡——就是测不准原理的那位;博尔赫斯——一位阿根廷的小说家,写过一堆很棒的短篇,现在因为跟 AI 关系密切而开始大火;还有康德。非常酷,简直炸裂,跟 AI 有很多有意思的交叠。强烈推荐。 >> 感觉光是你读的书、你推荐的书,我们就能做一整期播客。我知道这是你的一大热情所在。我现在着迷的是《The Power Broker》,我去你那儿做客时好像跟你聊过。它简直读不完,但读纽约的历史出乎意料地引人入胜。好,第二个问题。如果你有时间看电视的话,最近有没有特别喜欢的电影或剧? >> 我最近看了好多篮球,这算一个。我今年成了 Knicks 的球迷,所以那特别开心。另外我最近看了一部……应该算是迷你纪录剧集,叫《The Dark Wizard》,讲一个叫 Dean Potter 的人,他算是 Alex Honnold 成为 Alex Honnold 之前的 Alex Honnold。他有一种非常极端的性格,什么都徒手攀岩,然后又穿着翼装做 BASE 跳伞之类的,片子探讨了他的心理以及他后来的遭遇。我也说不清,我就是有点喜欢这类东西。还有一部叫《100 Foot Wave》,讲那些想冲大浪的人、大浪冲浪手。那种东西总有种……我猜它让我想起创业者吧。但《The Dark Wizard》,强烈推荐。 >> 有没有什么你最近发现、特别喜欢的产品? >> Codex。它简直是最好的,是真的好,真的好。 >> 你有没有一句最喜欢的人生格言,在工作或生活里经常会回到它? >> 有,还好几句。我大学时给自己写的那句核心的是:做值得被书写的事,写值得被阅读的东西。然后还有个叫 Rob Brezsny 的人,他在那种「AI 和冥想交叠」的圈子里——这也是个挺大的圈子——非常受欢迎。我也很喜欢他。他已经去世了,但我觉得他特别棒。我听了好多他的演讲,其中有一场,就一句话,他讲的是:当你在应对很难的事情时,你想要做到的,是能从一个开阔且有力量(spaciousness and strength)的位置去面对它。我觉得这里头有种特别有意思也特别重要的东西。很多冥想圈的论调,或者说一般人怎么应对难事,更偏向 David Goggins 那种——你就得硬上、就得豁出去那种,有时候那也确实管用。但我觉得有时候当你在应对一些事情时——比如说,我特别害怕 AI 会怎么改变我的工作——对我特别有帮助的是问自己:我是不是从一个开阔且有力量的视角在面对这件事?如果不是,我能不能让自己到达那个状态?因为从那个位置去处理它,对我来说会有效得多。这对我帮助非常非常大。 >> 哇,我太喜欢这句了。那我们最后一个问题,就扣着这场对话的主题:我好奇有没有哪个 AI 工具,你觉得现在还是被低估了的,你最近觉得……我是说——


[1:31:19] Lenny

are sleeping on. I I I I

……大家都没怎么重视的。我,我,我,我——


[1:31:21] Dan

Codex. I hate to say this but I have to because like any anyone who knows me like we were at this this conference recently at at an Anthropic conference and I'm like telling like Boris and Kat from Claude Code like you have to try Codex. And um it's it's just really good and the things that you can do with it are so different. Um uh especially if you're using it with the Anthropic browser to do things like your emails or check your check analytics or like anything like that. Um it it has completely transformed the way I work and I would be doing you a disservice if I like was searching for something else because it is that good. Damn, that's wild. Uh do you feel like Anthropic can catch up and or is this just like well No, yes, I think I think they can. I I like like I said, I think it's going to be a horse race and and different people will be ahead at different at different times, but I think right now Open AI has like has has gotten back the mandate of heaven a little bit. It's been It was a rough couple a couple months, like 6 months or so, but I think they're back. Interesting. And and you'd switch if one became more

Codex。我其实不太想这么说,但又不得不说,因为认识我的人都知道——前阵子我们参加了一场 Anthropic 的大会,我当时就跟 Claude Code 团队的 Boris 和 Kat 说:你们一定得试试 Codex。它真的非常好用,而且它能做的事情完全是另一个层次的。尤其是配合 Anthropic 的浏览器一起用,去处理你的邮件、查数据分析之类的活儿。它彻底改变了我的工作方式,如果我假装在找别的东西、不老实推荐它,那才是对你们不负责任,因为它就是那么好用。哇,这也太猛了。那你觉得 Anthropic 能追上来吗,还是说这就是……不,我觉得能,我觉得他们能追上来。就像我前面说的,我觉得这会是一场拉锯战,不同时间段会有不同的人领先,但我觉得现在 OpenAI 确实又拿回了一点「天命」。之前有一段时间挺艰难的,大概有六个月左右吧,但我觉得他们现在又支棱起来了。有意思。那如果哪一家变得更……你会换过去吗?


[1:32:28] Dan

I would. I would. People People It's funny. People are like, "Oh, are you like sponsored by Open AI?" And I'm like, "No, I just like talk about what I like." I was super loud about Claude Code when that was the thing I really liked. And I'll just say what I like when when it happens, you know? And to your point, people like there's a lot of value in using both for different things. So people There is. I I switch back and forth. Like I I truly do still use Claude a lot. Yeah. Such a big market. Well, Dan, we did it. We We went through so much. I can't wait to revisit this in a year {slash} get this out so people can start planning for this next year. Um two final questions. Where can folks find you and every what should people know and then how can listeners be useful to you? You can find me on X at Dan Shipper, s h i p p e r, and you can subscribe to every. Please subscribe to every, every.to. every.to/subscribe. How can listeners be useful? You know, have fun with AI. Like seriously, it's it's super fun. There's like a lot of It's not necessarily useful to me, but like it's it makes it I think it makes everything better when people put their hands in it and just like start figuring it out together rather than like arguing about it. And so the most useful thing you can do is like find ways to use it well in your life and share it. Dan, thank you so much for being here. Thank you. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennyspodcast.com. See you in the next episode.

我会,我会换。说来好笑,老有人问我:「你是不是收了 OpenAI 的钱?」我说:「没有啊,我就是单纯聊我喜欢的东西。」当年我对 Claude Code 也是疯狂安利,因为那时候它就是我最喜欢的工具。我就是喜欢什么就说什么,到时候自然就会说。而且就像你说的,针对不同的场景,两个一起用其实价值很大。确实是这样,我会来回切换。我是真的现在还经常用 Claude。是啊。这市场太大了。好了 Dan,我们成功收尾了,聊了好多东西。我已经迫不及待想一年后再回顾这期,或者赶紧把它发出去,好让大家提前为来年做规划。最后两个问题:大家可以在哪里找到你和 Every,有什么是大家应该知道的,以及听众可以怎样帮到你?你们可以在 X 上找到我,账号是 Dan Shipper,s-h-i-p-p-e-r,也欢迎订阅 Every。拜托大家订阅 Every,网址是 every.to,every.to/subscribe。听众能怎么帮我?我想说——尽情享受 AI 吧。说真的,它超级好玩。这其实不一定对我有什么直接的帮助,但我觉得当大家都真正上手去玩、一起摸索,而不是停留在争论上的时候,一切都会变得更好。所以你能做的最有用的事,就是去找到把它用好的方法,融入到自己的生活里,然后分享出去。Dan,非常感谢你来参加节目。谢谢你。也非常感谢各位的收听。如果你觉得这期有价值,可以在 Apple Podcasts、Spotify 或你喜欢的播客 App 上订阅本节目。也请考虑给我们打个分或留个评价,这真的能帮助更多听众发现这档播客。你可以在 lennyspodcast.com 找到所有往期节目,或了解更多节目信息。我们下期再见。