Complete Course: Claude for PMs (Cowork + Code + Dispatch)
频道: Aakash Gupta
视频: https://www.youtube.com/watch?v=bITUsUsrxjM
原文语言: en
统计: 共 63 轮 · Aakash 20 · Pavel 18
[0:00] Aakash
There's no good reason to talk to Claude just over a normal web chat anymore.
现在已经没有任何理由还只是通过普通的 web chat 去跟 Claude 对话了。
[0:03]
I don't use chat at all. I use it sporadically just if the window is open. I can ask some questions like is this grammatically correct?
我基本完全不用 chat。我只是偶尔用一下,比如窗口正好开着,我可以问几个问题,像是「这句话语法对不对?」
[0:10] Aakash
Claude is way better at using PowerPoint than it was. There is no excuse to walk into a meeting with a bad presentation anymore. You basically get a McKinsey level output in a minute or two. Coda is not adjusted to working with code base. As a product manager, you will be working with code bases a lot. Pavel Hern is the number one AIPM voice in Europe with over 200,000 LinkedIn followers and over 100,000 newsletter subscribers. And today he's going to break down everything you need to know to get the most out of Claude as a product manager.
Claude 现在用 PowerPoint 比以前强太多了。如今你再也没有借口带着一份糟糕的演示文稿走进会议室。基本上一两分钟就能拿到麦肯锡级别的产出。Claude 是专门为跟代码库打交道而调校过的。作为产品经理,你会经常跟代码库打交道。Pavel Hern 是欧洲排名第一的 AIPM 声音,在 LinkedIn 上有超过 20 万粉丝,newsletter 订阅用户超过 10 万。今天他要把你作为产品经理想最大化用好 Claude 所需要知道的一切,掰开揉碎讲清楚。
[0:39] Aakash
Pavel, you have spent more time in Coda than almost anyone. You use it for a lot of your everyday day-to-day tasks even though you were a former engineer who could use terminal just fine. So can you walk us through what a power user, what a master setup looks like in Coda and how you use it?
Pavel,你在 Claude 上花的时间几乎比任何人都多。你用它处理很多日常工作,尽管你以前是个工程师、用 terminal 完全没问题。所以能不能带我们看看,一个高级用户、一套大师级的配置在 Claude 里长什么样,以及你是怎么用它的?
[0:53] Pavel
In Coda, we can organize work in projects or folders. So what you can see here Most people stay in chat forever. That's like using Photoshop only to crop photos. And today we are going to discuss chat which is part of the cloud desktop, but also Coda, cloud code, dispatch which allows you to control remote sessions for Coda and code, web sessions, and how to use them all to 10x your productivity. Before we go any further, do me a favor and check that you are subscribed on YouTube and following on Apple and Spotify podcasts. And if you want to get access to amazing AI tools, check out my bundle. Where if you become an annual subscriber to my newsletter, you get a full year free of the paid plans of Maven, Arise, Relay App, Dovetail, Linear, Magic Patterns, Deep Sky, Reforge, Build, Descript, and Speechify. So be sure to check that out at bundle.akashg.com and now into today's episode. So Pavel is the guest behind our most popular episode ever with over 65,000 views when he gave an impromptu complete course on AI product management. He wrote did the NAN video which absolutely crushed with over 10,000 views. Before that, he did a discovery masterclass. Today, he is back for a record fourth time to help you understand how to get the most out of your Claude Enterprise subscription. When to use Claude on the web, when to use Claude co-work, when to use Claude code, when to use Claude dispatch. Pavel has tried all of the combinations of PM task versus tool and he's going to break down when you should use which tool for what. And he is going to break down how you can create a self-improving system that builds upon itself so that every time you do something, the system gets better at doing that task. Pavel, welcome back to the podcast. Hi Akash. Thanks for having me.
在 Claude 里,我们可以用项目或者文件夹来组织工作。所以你在这里能看到的——大多数人会永远停留在 chat 里。那就好比只用 Photoshop 来裁剪照片。而今天我们要聊的不只是 chat(它是 Claude 桌面端的一部分),还有 Cowork、Claude Code、Dispatch(它让你可以远程控制 Cowork 和 Code 的会话)、web 会话,以及如何把它们全部用起来,把你的生产力提高十倍。在我们继续之前,帮我个忙,确认一下你已经在 YouTube 上订阅、在 Apple 和 Spotify 播客上关注了。如果你想用上一堆超棒的 AI 工具,看看我的套餐包——只要你成为我 newsletter 的年度订阅用户,就能免费拿到一整年 Maven、Arise、Relay App、Dovetail、Linear、Magic Patterns、Deep Sky、Reforge、Build、Descript 和 Speechify 这些付费方案。所以一定去 bundle.akashg.com 看看。现在进入今天这一期。Pavel 是我们史上最受欢迎那一期的嘉宾,那期播放量超过 6.5 万——当时他即兴上了一整堂 AI 产品管理的完整课程。他还做了那期 NAN 视频,反响极佳,播放量超过 1 万。在那之前,他上过一堂 discovery 大师课。今天他第四次回到节目——创纪录——来帮你搞懂如何最大化用好你的 Claude 企业版订阅:什么时候用 web 上的 Claude、什么时候用 Claude Cowork、什么时候用 Claude Code、什么时候用 Claude Dispatch。Pavel 把「PM 任务」对「工具」的所有组合都试了个遍,他会讲清楚每种场景该用哪个工具。他还会讲清楚你怎么搭建一套能自我改进、不断累加的系统——这样每次你做一件事,系统就会越来越擅长做那件事。Pavel,欢迎回到播客。嗨 Aakash,谢谢你邀请我。
[2:51] Pavel
I love getting your newsletter in my inbox. One of the most interesting things you've been writing about recently is how Anthropic has gone from a billion to 30 billion in 16 months on the back of this insane shipping velocity. You tracked 74 releases in 52 days, which is pretty mind-blowing. Putting aside what's happening at Anthropic, what I want you to help us understand is what does this velocity say about the future direction of product management? What what I can see about serving Anthropic is that they are adjusting and many companies, not just Anthropic, are adjusting their workflows to AI so they do not use AI to replace steps in their processes, but to redesign their processes around what is currently possible. What I can see in many companies is that those roles, so product manager, product marketing manager, designer, and engineer are maybe not merging, but coming closer together. And many of the things that required a separate role, like maybe prototyping, maybe testing, writing release notes, even this designing an interface based on the design system, this is all being automated. And for a PM to be to thrive in this new environment and in those new workflows, product managers must understand technology. They must get comfortable with uh tools that previously were designed for engineers traditionally. So, like the terminal. Uh they also need to go outside their comfort zone and understand strategy, understand how they work drives revenue for the business, how it connects to to business goals, product strategy. Uh understand how uh it translate to to revenue. The future is super PM or super individual contributor uh with maybe a PM focus, maybe engineering focus. Uh but having skills from multiple areas, not just one. You cannot be a product
我很喜欢在收件箱里收到你的 newsletter。你最近写的最有意思的事情之一,就是 Anthropic 是如何凭着这种疯狂的出货速度,在 16 个月里从 10 亿做到 300 亿的。你统计到他们在 52 天里发布了 74 个版本,这相当令人震惊。先把 Anthropic 正在发生的事放一边,我想请你帮我们理解的是:这种速度对产品管理的未来方向意味着什么?我观察 Anthropic 能看到的是,他们在调整——而且不只是 Anthropic,很多公司都在把自己的工作流向 AI 调整。所以他们不是用 AI 去替换流程里的某些步骤,而是围绕「当下能做到什么」来重新设计整个流程。我在很多公司能看到的是,那些角色——产品经理、产品营销经理、设计师、工程师——也许不是在合并,但确实在彼此靠拢。很多以前需要单独一个角色来做的事,比如做原型、做测试、写发布说明,甚至基于设计系统去设计界面,现在全都在被自动化。要让一个 PM 在这种新环境、这些新工作流里如鱼得水,产品经理必须懂技术。他们必须对那些传统上为工程师设计的工具变得自如,比如 terminal。他们还得走出舒适区,去理解战略,理解自己的工作如何为业务带来收入,如何跟业务目标、产品战略挂钩,理解它如何转化成收入。未来属于「超级 PM」或者说「超级个人贡献者」——也许偏 PM、也许偏工程,但拥有跨多个领域的技能,而不只是单一领域。你不能只是一个
[5:03] Pavel
manager who only interviews customers and creates items in the product backlog. Yeah, the nature of the role is shifting, but for my money, this is the most fun version of product management. Yeah. Yeah, yeah, absolutely.
只会访谈客户、然后在产品 backlog 里创建条目的产品经理。是的,这个角色的性质正在转变,但在我看来,这是产品管理最有意思的一个版本。对,对,没错。
[5:16]
[laughter]
[笑声]
[5:16] Aakash
In the 16 plus years I've been in this field, I've seen it evolve a lot, which is one of its constants is that it's always evolving more than the average profession. And I think this version of it is the closest to the bare metal of the product, which I think is incredibly exciting. So, for this future model of product management, I want to break down for everybody who watches this episode to the end how to embrace it, how to actually redesign their processes from the ground up like Anthropic, How do you use the right tooling from Anthropic to behave like a next generation AI product manager. So, that's the goal for this episode today. We're going to break down for folks. You and I spend a good amount of time chatting on DMs and WhatsApp. And one of the craziest things you told me was that there is no good reason to use chat anymore. That is there's no good reason to talk to Claude just over a normal web chat anymore. Can you break this down because this is going to surprise a lot of product managers listening? Yeah, like I I personally I would lie if I told you that I don't use chat at all. I use it sporadically just if the window is open, I can ask some question like is this grammatically correct or something. But small most of the time when starting a session you don't know what exactly you will need. And when you start a session in chat, there are certain restrictions that you will face sooner or later. Like you cannot continue. So, let's imagine you have started your work, you're in the middle of the work, and you have to need uh leave your desktop. You cannot continue. You cannot copy the entire chat history to another window and maybe start a remote session, but you cannot continue in this in this window where where the conversations happened.
在我入行的这 16 年多里,我见证了它演变了很多——它的一个不变之处,就是它永远在演变,比一般的职业演变得更厉害。而我觉得它现在这个版本,是最贴近产品「裸机层」的,我觉得这相当令人兴奋。所以,针对产品管理的这个未来形态,我想为每一个把这期节目看到最后的人拆解清楚:怎么去拥抱它,怎么真正像 Anthropic 那样从头重新设计你的流程,怎么用好 Anthropic 的工具,表现得像一个新一代的 AI 产品经理。这就是今天这期的目标。我们会为大家拆解。你和我在 DM 和 WhatsApp 上聊了不少。你跟我说过的最疯狂的一件事,就是现在已经没有任何好理由再用 chat 了——也就是说,已经没有任何好理由还只是通过普通的 web chat 去跟 Claude 对话了。能不能展开讲讲?因为这会让很多正在听的产品经理大吃一惊。是的,我个人——如果我跟你说我完全不用 chat,那是撒谎。我会偶尔用一下,就是窗口正好开着的时候,我可以问个问题,比如「这句话语法对不对」之类的。但大多数时候,当你开始一段会话时,你并不知道自己具体会需要什么。而当你在 chat 里开始一段会话,你迟早会撞上一些限制。比如你没法继续。我们设想一下:你已经开始工作了,工作进行到一半,然后你得离开电脑了。你没法继续。你没法把整段聊天记录复制到另一个窗口,然后也许开一个远程会话,但你没法在原来发生对话的这个窗口里继续。
[7:05] Pavel
Uh you cannot continue on mobile. Um there are restrictions related to what if you decide that now I want to code something. Uh chat cannot cannot do that. You want to create an HTML page then export this HTML page to infographic, and you need this infographic in your email. Uh so, once again, you need to start a different session with a different context and explain what you have been discussing with chat. So, for me it's easier just to start everything either in with code work with dispatcher or with cloud code, and then you can easily navigate between those interfaces. Does it make sense, Akash? Yeah, it does. So, I've been building a lot of AI products lately. My job search OS has 16 different agents. My newsletter has a recommendation engine, and I kept running into the same problem. I'd ship something, it would work in my testing, and then I'd get messages from users saying it's hallucinating or picking the wrong tool. The issue wasn't the prompts or the tools, it was that I wasn't actually evaluating anything. I didn't have a way to see what my agent was actually doing step by step. Every tool call, every decision. That's where Arize comes in. Let me show you. I'm going to open Claude Code and install Arize with just one command. NPX skills add arizeai arize skills skill yes. Now, Claude Code already knows how to instrument my agent. I tell it set up tracing to Arize, and it automatically analyzes my code base, figures out where the LLM and tool calls are, and adds instrumentation automatically. Now, I can see everything, every trace, every span, every decision. And more importantly, I can evaluate it. That's the shift. Trace what's happening, evaluate where it fails, then fix it. This trace right here, my resume feedback agent was
你也没法在手机上继续。还有一些限制,比如:万一你现在决定「我想写点代码」,chat 做不到。又比如你想做一个 HTML 页面,然后把这个 HTML 页面导出成信息图,而你需要把这张信息图放进你的邮件里。所以你又得开一段不同的会话、带着不同的 context,再把你刚才跟 chat 讨论的内容重新解释一遍。所以对我来说,更省事的做法是:一切都直接在 Cowork、Dispatch 或者 Claude Code 里开始,然后你就能轻松地在这些界面之间来回切换。这样讲说得通吗,Aakash?说得通。所以,我最近一直在做很多 AI 产品。我的求职 OS 里有 16 个不同的 agent。我的 newsletter 有一个推荐引擎,而我一直撞上同一个问题:我上线一个东西,在我自己测试时它好好的,然后我就会收到用户来信说它在胡编、或者选错了工具。问题不在 prompt 或工具上,而在于我其实根本没在评估任何东西。我没有办法看到我的 agent 到底一步一步在做什么——每一次工具调用、每一个决策。这正是 Arize 派上用场的地方。我给你看看。我打开 Claude Code,用一条命令就装好 Arize:NPX skills add arizeai arize skills skill yes。现在 Claude Code 已经知道怎么给我的 agent 加埋点了。我跟它说「把 tracing 接到 Arize」,它就会自动分析我的代码库,找出 LLM 和工具调用在哪儿,然后自动加上埋点。现在我就能看到一切——每一条 trace、每一个 span、每一个决策。更重要的是,我可以对它做评估。这就是关键转变:trace 发生了什么,evaluate 在哪儿失败,然后 fix。就拿眼前这条 trace 来说,我的简历反馈 agent 本来
[8:53] Aakash
supposed to pull the company's tech stack from the job posting, but instead, it hallucinated that they use React when the posting said Python. But instead, it hallucinated that they use React when the posting said Python. I never would have caught that without seeing the trace. And here's the part that blew my mind. I asked Claude Code to look at these traces and tell me what I should be evaluating. It came back with four eval criteria I hadn't written. Things like picking the right tool and staying grounded in the input. I wrote the evals, ran them, and found that my agent was making the same kind of mistake about 12% of the time. Claude pushed a fix, I re-ran the evals, and it dropped to under 2%. That whole loop, trace, evaluate, fix, took me about 20 minutes, and now it runs automatically. If you're building AI products and not evaluating them, you're shipping blind. Try Evolve free at evolve.com and get a year free, a $1,260 value with my bundle. Evolve, check it out. It's one of the top AI Evolve platforms used by all of the top AI teams for a reason. Let's make this more concrete though. Most PMs I talk to, they're finally on like a Claude Pro or Max or API Enterprise subscription. What they really need is a very clear mapping of if this, then that. You've mapped all of this out for us. Can you show us when should a product manager be using Co-work versus Dispatch versus Code? So, Chat is like it's like a chatbot like ChatGPT with various tools. So, you type a question, it answers. Sometimes it can execute a simple script and for example, reply with a spreadsheet, but other than that uh nothing complex. So, it can draft an email, it can pretend to be some person, summarize information, very basic stuff. Uh when it comes to Co-work, it's about
应该从招聘启事里抓出公司的技术栈,但它却胡编说他们用 React——而招聘启事里写的是 Python。它却胡编说他们用 React,而招聘启事里明明写的是 Python。要不是看到这条 trace,我根本不可能发现这个问题。而最让我震撼的是这一部分:我让 Claude Code 去看这些 trace,告诉我应该评估些什么。它给我提了四条我自己都没写出来的评估标准,比如「选对工具」和「紧扣输入、不脱离原文」。我把这些 eval 写好、跑起来,发现我的 agent 大约有 12% 的概率在犯同一类错误。Claude 推送了一个修复,我重新跑了一遍 eval,错误率降到了 2% 以下。整个这个循环——trace、evaluate、fix——总共花了我大概 20 分钟,而现在它会自动运行。如果你在做 AI 产品却不去评估它,那你就是在盲发。去 evolve.com 免费试用 Evolve,用我的套餐包还能免费拿一年,价值 1260 美元。Evolve,去看看吧。它是顶尖的 AI 评估平台之一,所有顶级 AI 团队都在用,这是有原因的。不过咱们把这件事说得更具体些。我接触到的大多数 PM,他们终于用上了 Claude Pro 或 Max 或者 API 企业版订阅。他们真正需要的是一个非常清晰的「如果……那就……」的对应关系。这套东西你都给我们梳理好了。能不能给我们演示一下:产品经理什么时候该用 Cowork、什么时候用 Dispatch、什么时候用 Code?好的,Chat 就像——它就像一个带各种工具的聊天机器人,类似 ChatGPT。你打一个问题,它回答你。有时候它能执行一个简单的脚本,比如回你一个表格,但除此之外就没什么复杂的了。它能起草一封邮件,能扮演某个人物,能总结信息,都是很基础的东西。而说到 Cowork,它讲的是
[10:46] Pavel
working with real files and executing workflows. So, by real files, I mean uh reorganize invoices on my desktop or create HTML infographic. Uh Co-work can also plan uh long-running tasks. So, for example, if completing a task requires taking five, six, seven, or more steps, it can do that and it can then execute those steps one by one. It can also spawn uh sub-agents. So, some tasks can be executed in parallel. So, for example, you want to write an email, you want to summarize your product strategy. Inside this email, you also want to add some presentation as an attachment, but the presentation should be converted to PDF. So, it can start one agent to to summarize strategy or whatever, another agent to create a HTML infographic, and then maybe convert it to PDF, and then another one to do something else. Maybe find personal details in your in some CRM system. And then, all those agent agents finish the work, it can get the results, maybe replay some steps or adjust, and finish the work. So, those are real real workflows, real real identical workflows with the access to real files on your desktop. And code is is the same, but it's coding. So, except other than connecting to all those systems and working with real files, it can execute scripts on your real machine because Co-work runs in a virtual machine, and code can execute scripts, can execute system commands on your laptop. And it can also it is also adjusted to work with code bases. So, it means it has different set of plugins, not adjusted to knowledge work, but adjusted to designing front end, working with databases, debugging, and so on. A PM who's non-technical who's watching this is probably thinking code is scary, the IDE is scary. Can I just skip it? No, you can't because engineers are going to use code.
处理真实文件、执行工作流。我说的「真实文件」,是指比如「把我桌面上的发票重新整理一下」或者「做一张 HTML 信息图」。Cowork 还能规划长时间运行的任务。比如说,如果完成一个任务需要走五步、六步、七步甚至更多步,它能做到,然后它会把这些步骤一个接一个地执行。它还能派生出 subagent,所以有些任务可以并行执行。举个例子:你想写一封邮件,你想总结你的产品战略,在这封邮件里你还想加一份演示文稿作为附件,但这份演示文稿要先转成 PDF。那它就可以启动一个 agent 去总结战略之类的,另一个 agent 去做一张 HTML 信息图、然后也许把它转成 PDF,再来一个 agent 去做别的事——比如在你某个 CRM 系统里找到一些个人资料。然后等所有这些 agent 都干完活,它就能把结果收回来,也许重放某几步、或者做些调整,最后把活儿收尾。所以这些是真正的工作流,真正、不折不扣的工作流,而且能访问你桌面上的真实文件。Code 也是一样的,只不过它是在写代码。所以除了连接所有这些系统、处理真实文件之外,它还能在你真实的机器上执行脚本——因为 Cowork 跑在一个虚拟机里,而 Code 能执行脚本、能在你的笔记本上执行系统命令。而且它也是专门为跟代码库打交道而调校过的。也就是说它有一套不一样的插件,不是为知识工作调校的,而是为做前端、操作数据库、调试等等调校的。一个非技术背景、正在看这期节目的 PM,心里大概在想:Code 看着吓人,IDE 看着吓人,我能不能直接跳过它?不行,你不能,因为工程师会用 Code。
[13:09]
[gasps]
[倒吸一口气]
[13:09] Pavel
And even though many features when talking about personal productivity, like working with files, analyzing information, even organizing, self-improving knowledge database, you can do that in Co-work. You will not have the explorer review, which is about uh, presenting the hierarchy of folders and files, but other than that, uh, Coda can do almost everything. Uh, but there are certain features in Coda that only Coda has. So, for example, uh, Coda has So, for example, one one example can be sub agents that you define in your solution. Coda cannot do that. It can call dynamic sub agents, but you you cannot have those agentic personas that you So, for example, the researcher, for example, an agent that tests your solution, an agent that, uh, creates release notes. You you cannot have this structure in Coda. And similarly, there are also other features that are like Coda specific or working with Coda bases specific, like uh, hooks that do not work in Coda. So, eventually, as a product manager working with engineers, you will have to work with Coda. I can't emphasize this enough. As a product manager, you should be learning Coda. You need to get over the initial sort of hump of this doesn't look great. But, Coda Coda is super powerful. And so, Pavel, you have spent more time in Coda than almost anyone. You use it for a lot of your everyday day-to-day tasks, even though you were a former engineer who could use terminal just fine. So, can you walk us through what a power user, what a master setup looks like in Coda, and how you use it? Mhm. Okay. So, in Coda, So, first thing to understand that, uh, we can organize work in projects or folders. Uh, so, what you can see here, here are some of the recent folders that I have been worked in working in, and also
而且,虽然在谈到个人生产力时,很多功能——比如处理文件、分析信息、甚至整理一个能自我改进的知识库——你都可以在 Cowork 里做。你不会有那个文件浏览器视图(也就是把文件夹和文件的层级展示出来),但除此之外,Claude 几乎什么都能做。不过有些功能是只有 Claude Code 才有的。比如说,举个例子,你可以在自己的解决方案里定义 subagent,这个 Cowork 做不到。它能调用动态 subagent,但你没法拥有那种你自己定义好的「agent 人设」。比如说,研究员;比如说,一个负责测试你方案的 agent;一个负责写发布说明的 agent。这种结构你在 Cowork 里没法搭。同样地,还有一些功能是 Claude Code 特有的、或者说是跟操作代码库特有的,比如 hook,它在 Cowork 里不起作用。所以归根到底,作为一个跟工程师协作的产品经理,你终归要用 Claude Code。我怎么强调都不为过:作为产品经理,你应该去学 Claude Code。你得先迈过「这看起来不太友好」这道坎。但 Claude Code 真的超级强大。所以,Pavel,你在 Claude 上花的时间几乎比任何人都多。你用它处理很多日常工作,尽管你以前是个工程师、用 terminal 完全没问题。所以能不能带我们看看,一个高级用户、一套大师级的配置在 Claude 里长什么样,以及你是怎么用它的?嗯。好的。在 Claude 里,首先要搞懂的是:我们可以用项目或者文件夹来组织工作。所以你在这里看到的——这些是我最近一直在用的一些文件夹,还有一些
[15:21] Pavel
predefined projects that can have pre-defined custom instructions. So, one of those projects is editor, which is my project that helps me with copywriting, research, generating infographics, and other things related to content creation. Uh but maybe we can first use it without it. So, I will just select some random folder on my desktop. It's not entirely random because I prepared it before.
预先定义好的项目,它们可以带有预先设好的自定义指令。其中一个项目叫 editor,这是我用来帮我做文案、做研究、生成信息图,以及其他跟内容创作相关的事情的项目。不过也许我们可以先不用它。我就随便选一个我桌面上的文件夹。其实也不是完全随便选的,因为我事先准备过。
[15:55]
[laughter]
[笑声]
[15:57] Pavel
Okay. So, here is the folder with invoices. Uh and yeah, after granting permissions, Cork can now access real files uh in this folder. And the files inside are just random invoices that I collected from 2 months. So, like that Atlassian invoices. I'm not sure if previous relevant, but yeah, like different invoices, uh different uh receipts and so on. Uh so, what I can now ask it to do is to organize, analyze PDF invoices in in this folder and group them by month in folders like uh Jan, Jan, Mar, Jun, etc. Uh yeah. The duplicate the duplicate files. Some of the files are duplicates, so which it should figure out how to do that. The names might be different, but the content is the same. And so now you guys are seeing immediately why co-work pretty much always is preferable to chat. In this case, it's actually manipulating things in your file system. Yeah. And as we can see, it created a list of steps, so it needs to extract dates from PDF invoices. It needs to identify and remove duplicates. Maybe it will use a hash function. It needs to create month folders and move files and yeah, verify that if everything is correct. Uh what we can see is also the context. So, context for this um because this is an agent, the context for this agent is um this file cloud.md instructions. Um There are no special instructions inside, but if there if this file existed, that would be this cloud.md that everyone is talking about with like custom instructions when working in this folder. Uh so, the next time I could prepare an inbox and just drag and drop new invoices inside and every time a new file appears, uh this process is repeated and the file is sent to the right subfolder. And it also dynamically loaded skills, which are like procedures. Uh skills are activated based on the task that
好。这是一个装着发票的文件夹。然后,对,在授予权限之后,Cowork 现在就能访问这个文件夹里的真实文件了。里面的文件就是我攒了两个月的一些随机发票。比如 Atlassian 的发票。我不确定之前那些是不是还相关,但反正就是各种各样的发票、各种各样的收据等等。那我现在可以让它做的,就是「整理、分析这个文件夹里的 PDF 发票,并按月份分到文件夹里」,比如 Jan、Jan、Mar、Jun 等等。对。还有重复文件——有些文件是重复的,它得自己想办法处理这个。文件名可能不一样,但内容是相同的。那现在大家立刻就能看出来,为什么 Cowork 基本上总是比 chat 更可取。在这个例子里,它实际上是在操作你文件系统里的东西。对。我们可以看到,它列出了一串步骤:它需要从 PDF 发票里提取日期,需要识别并删除重复项——也许它会用一个哈希函数,需要创建月份文件夹、移动文件,最后,对,还要验证一切是否正确。我们还能看到的是 context。这个 agent 的 context,就是这个 CLAUDE.md 指令文件。里面没有什么特别的指令,但如果这个文件存在,那它就会是大家都在谈的那个 CLAUDE.md——里面写着「在这个文件夹里工作时」的自定义指令。所以下一次,我可以准备一个 inbox 文件夹,直接把新的发票拖进去,每当出现一个新文件,这个流程就会重复一遍,文件就会被送到对应的子文件夹里。它还动态加载了 skill——skill 就像一套操作流程。skill 是根据 agent 当前正在执行的任务而激活的。
[18:52] Pavel
the agent currently is executing. So, skill has a description and based on this description, an agent can decide that yeah, this skill is about working with PDFs. I'm working with PDFs, so let's see what is inside and then it will read detailed instructions, detailed procedures or of how to work with PDF files. And and we call this progressive disclosure. So, you can have dozens or of hundreds maybe of skills. And Cloud will read them. So, Cloud will read the detailed instructions only when the skill description matches what you are trying to do. So, in this case only one PDF was identified. Uh what it did, it created four folders. You can see it. It removed some duplicates. And the names suggest that suggest that indeed those were duplicates. Let's see the folder. Okay, April, February, Jan, March. If I open January, then I have some invoice from January. It's This is not a secret. Maybe this This one is. So, let's try a different one. Um yeah, like post mark is it is not a secret. I can present it. So, it's February. Is that correct? Yeah, it's it's an invoice from February. And that's it. And I asked about PDF, but you can also see that it also processed images. Yeah. So, yeah, and this is some fuel invoice. Uh let me verify that it is correct. It's in Polish, but yeah, it's February. I hope you can see it. So, it just ident- it just understood that maybe Pavel didn't know that there are images inside. Let's Let's move them to this folder, too. Uh but
所以一个 skill 会有一段描述,agent 根据这段描述就能判断:嗯,这个 skill 是讲怎么处理 PDF 的,而我正在处理 PDF,那就看看里面有什么——然后它就会去读详细指令,读处理 PDF 文件的详细操作流程。我们把这叫做渐进式披露(progressive disclosure)。所以你可以有几十个、甚至上百个 skill,而 Claude 会去读它们——Claude 只会在某个 skill 的描述跟你当前想做的事情匹配时,才去读它的详细指令。所以在这个例子里,只识别出了一个 PDF。它做的是:创建了四个文件夹,你能看到;它删掉了一些重复项,而文件名也表明那些确实是重复的。我们看看文件夹。好的,April、February、Jan、March。我打开 January,里面有一张一月的发票。这张……这个不是机密。也许这张……这张是机密的。那我们换一张。嗯,对,比如 Postmark 这张就不是机密,我可以展示出来。这是二月的。对不对?对,这是一张二月的发票。就是这样。我问的是 PDF,但你也能看到它还处理了图片。对。所以,对,这是一张加油的发票。我来核对一下它对不对。它是波兰语的,但对,是二月的。希望你能看清。所以它就……它就只是理解到,也许 Pavel 自己都不知道里面还有图片,那「我们也把它们移到这个文件夹里吧」。不过……
[20:59] Aakash
Not Gemini is not the only thing that can read images these days. Cloud can, too. Yeah, no problem. What else it can do? Um so, it can work with files, so we have presented that. It can load skills, which are instructions, but skill can also have um scripts to be executed inside, not necessarily by Yeah, Co-worker can also execute them inside a virtual machine. Uh it can connect to external and local services, and the most popular uh format is MCP server. So, it's like the USB for for agents. And by MCP servers, I mean in cloud, they are called called connectors. So, we have connector for Google's for Google Drive, for Gmail, for Slack, and for dozens of other apps or hundreds of other apps. Uh some of those connectors of MC all the MCP servers are provided by Antropic and others. You can just download them and configure locally. And yeah, talk to your apps. So, for example, I can ask Let me demonstrate. Um We are still in this folder. It doesn't matter, but uh How many un- unanswered emails I have? Uh this is a good one.
现在能读图的可不止 Gemini,Claude 也行。对,没问题。它还能干什么?嗯,它能处理文件,这个我们刚才演示过了。它能加载 skill,skill 本质上是一组指令,但里面也可以带要执行的脚本,执行的不一定非得是——对,Cowork 也能在虚拟机里把脚本跑起来。它还能连外部和本地的服务,最流行的格式就是 MCP server,相当于 agent 的 USB 接口。说到 MCP server,在 Claude 里它们叫 connector(连接器)。所以我们有连 Google Drive、Gmail、Slack 的 connector,还有几十上百个其他应用的。其中一部分 MCP connector 是 Anthropic 提供的,还有别人做的,你直接下载、在本地配置好,就能跟你的各种应用对话了。比如说我可以问——我来演示一下。嗯,我们现在还在这个文件夹里,无所谓——我有多少封没回的邮件?嗯,这个问题问得好。
[22:30]
[laughter]
[笑声]
[22:31]
I don't want to run this on my inbox live on the podcast. Do not reveal any personal information or emails. Yeah, let's let us see. Uh Uh so, it connected to my Gmail account. It can also draft emails, and it can This one cannot send. It can only create the drafts, but you can also connect a connector that can can send emails for you. How do you process your email inbox today? Do you recommend most PMs do it with Cloud? Yeah, I create all the responses. Yeah, this is my current configuration. So, yeah, Cloud the system that learns drafts replies. Um but it doesn't send them automatically. So, I'm using this default connector that cannot send. And similarly on Slack, theoretically, you can um ask Cloud to respond automatically if you use this connector. Uh there is a footer sent by Cloud, so everyone can see it. Uh instead, I ask it to draft replies and I verify them. Uh in many cases, someone asks about the URL, so no reason to find for this Yeah, to to to spend time looking for this and Uh I just approve the messages or edit them. I send it manually, so there is a button in the interface, send. And after every session, Cowork or Code, depending on what I what interface I'm using, uh verifies my responses and tries to learn from them. So, the next time it will get better. Love it. Uh and that's basically that's basically it. So, so a big a big thing are those skills that you can you can either use uh predefined skills or skills from third-party marketplaces like like my micro marketplace. Um that's basically it when it comes to Cowork. So, Can you show us the PM skills marketplace? This hit 1,300 GitHub stars in 72 hours. So, it's gone pretty viral. Yes, of course. So, it's this one. It's p h u r y n p m skills / p m skills. It's currently
我可不想在播客上直接拿我的真实收件箱来跑。「不要泄露任何个人信息或邮件内容。」好,咱们来看看。嗯,它连上了我的 Gmail 账号。它也能起草邮件,不过这个 connector 不能发送,只能创建草稿,但你也可以接一个能替你发邮件的 connector。你现在每天怎么处理收件箱?你会推荐大多数 PM 都用 Claude 来处理吗?会,所有回复都是我让它生成的。对,这就是我目前的配置。所以 Claude 这套系统会学习、起草回复,但不会自动发送。我用的是这个默认的、不能发送的 connector。Slack 上同理,理论上你可以让 Claude 用这个 connector 自动回复,不过会有一行「由 Claude 发送」的页脚,大家都看得到。我没这么干,而是让它起草回复,由我来核对。很多时候有人就是来问个网址,没必要专门花时间去找——对,没必要去翻这个——我就直接批准这些消息,或者改一改。然后手动发出去,界面里有个「发送」按钮。每次会话结束后,Cowork 或者 Code(看我用的是哪个界面)会复核我的回复,从中学习,下次就做得更好了。太棒了。嗯,基本上就是这些。所以一个很大的亮点就是这些 skill,你可以用预置的 skill,也可以用第三方市场里的 skill,比如我那个小市场。Cowork 这块大致就是这些。那能给我们看看那个 PM skill 市场吗?它 72 小时就拿到了 1300 个 GitHub star,挺火的。当然可以。就是这个,phuryn/pm-skills,路径是 pm-skills,现在已经
[25:06] Aakash
10,000 GitHub stars and what I have done is that I I I created a set of uh plugins. So, a plugin is like a collection of skills and commands and for different domains like data analytics, go to market, market research and you can upload each of those plugins separately. Next, inside each plugin you have skills. So, for example, if we open, let's say product discovery, inside there are Ah, there's a documentation. So,
一万个 GitHub star 了。我做的事情是:我创建了一组 plugin(插件)。一个 plugin 就是一堆 skill 和命令的合集,对应不同领域,比如数据分析、go-to-market(市场推广)、市场调研,这些 plugin 你可以单独上传、单独安装。然后每个 plugin 里面又有一堆 skill。比如我们打开「product discovery(产品探索)」这个,里面有——啊,这里有份说明文档。所以
[25:45]
[laughter]
[笑声]
[25:46] Aakash
there are analyze feature requests, brainstorm ideas, plan experiments, um create metrics to track um your feature and so on. And also, I have defined uh workflows that aggregate more than one skill. So, for example, product discovery, this can be um Okay, this is a a better example. So, discover, it is from ideation, then we map assumptions, then we think about uh we think about uh so, we analyze Actually, this description is not correct. So, it analyzes customer needs, then based on those needs, it's uh will map the opportunities and how important certain problems are for the customers and how satisfied they are with what they already have. Then, it will ideate, so how we can solve those problems, uh map the assumptions related to value, usability, feasibility, viability, maybe some ethical considerations and it will plan plan experiments to prove or disprove our assumptions. Uh so this is like the entire product product discovery workflow in one comment. Uh how to use it? You can go to the home page and there is a a script, but basically you can just take this URL go to cloud uh open customize. Let me see. Personal plugins. Add marketplace. Yeah, like uh Yeah, this interface is really changing all the time. So uh it's uh let's see this one, but I think that owner {slash} repo doesn't work despite the description. So let's let's try this one. That's my gut feeling that this this can work. Okay, and uh Those are not mine. So yeah, we can see it here. So data analytics execution, go-to-market, market research, and so on. And if I click product discovery or product strategy. Yeah, maybe this one. Now it is added, installed, and ready to use. What is inside? Analyze. Ah. Strategy. We have Ansoff matrix, pricing strategy. So now when I go to to the chat uh
里面有「分析功能需求」「头脑风暴」「规划实验」、嗯、「为你的功能设定追踪指标」等等。另外我还定义了一些 workflow(工作流),把不止一个 skill 串起来。比如产品探索这个,可以是——好,这个例子更合适——「discover(探索)」,它从构思开始,然后梳理假设,再去思考——嗯,我们先去分析——其实这段描述不太对。它会先分析客户需求,再根据这些需求映射出机会点,看看某些问题对客户有多重要、客户对现有方案有多满意。然后开始构思,也就是我们怎么解决这些问题,再梳理跟价值、可用性、可行性、商业可行性相关的假设,可能还有一些伦理上的考量,最后规划实验去验证或推翻这些假设。所以这一条命令里,就是一整套产品探索的工作流。怎么用?你可以去主页,那里有个脚本,但基本上你只要拿到这个 URL,去 Claude 里打开「自定义」——我看看——「个人插件」,「添加市场」。对,就像——对,这个界面一直在变。嗯,那个——我们看看这个,不过我觉得描述里说的 owner/repo 这种写法好像不管用。咱们试试这个,我直觉觉得这个能行。好,嗯,那些不是我的。对,我们能在这看到:数据分析、执行、go-to-market、市场调研等等。如果我点「产品探索」或者「产品战略」——对,可能这个——现在它就添加好、安装好、可以用了。里面有什么?「分析」。啊,「战略」。我们有 Ansoff 矩阵、定价战略。那现在我去聊天界面,嗯
[28:38] Aakash
Let's start a new task in co-work. Help me design And theoretically I can count on skill being activated automatically, but yeah, if you want to really be sure that the skill is activated. Uh yeah. Use the slash command, which I do recommend. Yeah, it's it's kind of messy and sometimes, if Cloud has general knowledge about a certain area, like uh product strategy, it thinks it knows something, but you want to override this knowledge and without doing it explicitly, it can default to to the training data. So, yeah, if you know that there is for Amazon 2.0, whatever it means. So, it loaded the skill and now it will interview me to get more information, but it is it is all part of the skill. So, what is the core concepts? Output format and then it will create product product strategy canvas, by the way, based on my product strategy canvas. When you're writing Google product strategy canvas, it will be mine. Uh the one I created. Nice. So, it looks like this and of course, there are those famous plugins that Anthropic created. The GitHub repo with plugins and skills. The main plugin was like legal. Uh hundreds of millions of dollars evaporated from uh from the stock market. Just because people understood how easy it is to to describe certain processes and automate them. Yeah, with a proper markdown file written by a subject matter expert, these days, Opus 4.5 and Opus 4.6, they can really use the tools, whether it's PowerPoint or Excel, or for PMs, Notion or Google Docs. And they can operate it like a competent person if they have the right instructions. It's pretty crazy. So, can you walk us through a couple of these skill files and how you constructed them and how somebody constructs a good con- skill file? Yeah, sure. Uh so, the skill is basically how to do
在 Cowork 里开一个新任务,「帮我设计——」理论上我可以指望这个 skill 被自动激活,不过,对,如果你想真正确保 skill 被激活,嗯,对,就用斜杠命令,这个我强烈推荐。因为有时候 Claude 对某个领域有通用知识,比如产品战略,它觉得自己懂,但你想覆盖掉它这部分知识,如果不明确指出来,它就可能退回去用训练数据里的东西。所以,对,如果你知道有这么个 skill——比如「Amazon 2.0」,不管这是什么意思。它现在加载了这个 skill,接下来它会采访我来获取更多信息,但这整个过程都是 skill 的一部分。所以「核心概念是什么?」「输出格式」,然后它会创建产品战略画布——顺便说一句,是基于我的产品战略画布。所以当你写「Google 产品战略画布」时,出来的会是我那套,是我创建的那一版。不错。它长这样,当然还有 Anthropic 出的那些有名的 plugin。那个装着各种 plugin 和 skill 的 GitHub repo,最主要的那个 plugin 当时算是个传奇了——好几亿美元的市值从股市上蒸发了,就因为人们意识到把某些流程描述出来、自动化掉是多么容易。对,只要有一份由领域专家写好的、规范的 markdown 文件,现在 Opus 4.5、Opus 4.6 就能真正地用好这些工具,无论是 PowerPoint、Excel,还是对 PM 来说的 Notion、Google Docs。只要给对了指令,它们就能像个称职的人一样操作这些工具,挺疯狂的。那你能不能带我们过几个 skill 文件,讲讲你是怎么构建的,以及别人该怎么写出一份好的 skill 文件?当然可以。嗯,skill 本质上就是「怎么做
[31:03] Aakash
something. So, some procedure or a domain knowledge about doing something. Uh So, you you don't really have to create those skills manually. You can describe this process to Co-worker Cloud Code and ask it to create a skill. So, so, basically, if you can teach a graduate how to do something, you just repeat the same to Cloud and it will create a skill. That's That's the easiest the easiest way. We can, of course, look inside and uh I usually don't do that. So, I just use the chat interface to to talk to Cloud. But, yeah, uh it's not this repo, it's this one. So, if we open, for example, product discovery skill, uh identify assumptions for an existing product. We have this skill.md file. Mm. And you can also have additional files in this folder. But, yeah, basically, the main file is skill.md. And And the format is markdown. Uh we can see the preview in GitHub, but if I switch to the code view, there is this intros. Mm. There is this section that uh agents read. So, they do not load the the skill. They load um name and description and when to use this skill, what it is for, and as I said, when doing a specific work like uh use when stress testing a future idea. So, if the user will s- write in the chat that, "Hey, I want to test this idea and uh assess the risk." the agent will see this skill. It will description and it it will just um trigger it. It will use it. Uh other than that, this is just a prompt. So, it can have instructions. It can be uh general knowledge without step-by-step, but it can also be like step one, uh get information from place A. Step B is format it in a specific way. Step Step three is something. So, it's it's just a prompt. So, in this case, it is uh yeah. It's the context. So, you are a devil's advocate. We are stress testing idea and there are
某件事」。就是某个流程,或者关于怎么做某件事的领域知识。嗯,你其实不用手动去创建这些 skill。你可以把这个流程描述给 Cowork、Claude Code,让它替你创建一个 skill。基本上,只要你能教会一个应届毕业生怎么做某件事,你把同样的话讲给 Claude,它就能创建出一个 skill。这是最简单的办法。我们当然也可以打开看看里面——我一般不这么做,我就用聊天界面跟 Claude 对话。不过,对,嗯,不是这个 repo,是这个。如果我们打开,比如说产品探索的 skill,「为一个现有产品识别假设」,我们有这个 skill.md 文件。嗯。这个文件夹里你还可以放别的附加文件,但基本上主文件就是 skill.md,格式是 markdown。我们能在 GitHub 里看预览,但如果我切到代码视图,会看到开头这部分。嗯。有这么一段是 agent 会读的,它们不会加载整个 skill,只加载名称和描述、什么时候该用这个 skill、它是干嘛的,就像我说的,在做某类具体工作的时候——比如「在压力测试一个未来想法时使用」。所以如果用户在聊天里写「嘿,我想测试这个想法、评估一下风险」,agent 就会看到这个 skill,看到它的描述,然后就会触发它、用上它。除此之外,剩下的就是一个 prompt 了。它可以是指令,可以是没有分步骤的通用知识,但也可以是「第一步,从 A 处获取信息;第二步,按特定格式整理;第三步,做某件事」。所以它就是个 prompt。具体到这个例子,它是,对,是 context(上下文)。「你是一个唱反调的人,我们在压力测试一个想法,以下是这些
[33:45] Aakash
those arguments. So, you we have seen that cloud uh when working on strategy asked me a few questions. So, those are the arguments. And then instructions, what it should do step-by-step. And that's all. And And by the way, this is uh if someone wants to learn more, they can uh either they can see it in as a markdown file or if they keep talking to an agent, uh the agent will suggest my articles about identifying assumptions. So, this is like like marketing inside skills, but relevant to what the person was trying to do. So, all the knowledge is inside skill, but then the agent will also have those articles uh a short-term memory. I would tell everybody like iterating on your skills is one of the highest ROI activities I personally have done. So take Pavel's skills as a baseline so at least you have a baseline. But then as you encounter some feedback for his skill, give that feedback to Claude and say, "I want you to improve my assumption existing skill." Read our chat and see the feedback I gave you, understand the root cause of what drove you to give a output that I had to give feedback. And rewrite the skill from first principles so that it doesn't make that mistake again. I have found this to be the single highest ROI activity I do so I think once you get the initial skill, you got to really iterate on it to really get the ROI. Like I agree, Akash. I do not memorize prompts and it's like in evals and you also have been writing about evals that you need to when building some AI system on some AI pipeline, you need to see how the system performs in real life and then identify failure modes. And in this case you don't create evals but you can just give feedback to to Claude, say what was wrong. It will understand the context because it already has this
论点。」所以——我们刚才看到 Claude 在做战略的时候问了我几个问题,那些就是这里的论点。然后是指令,告诉它该一步步做什么,就这样。对了,顺便说一句,如果有人想了解更多,他们要么把它当 markdown 文件来看,要么如果他们继续跟 agent 对话,agent 会推荐我那些关于「识别假设」的文章。这就像是 skill 内部的营销,但又跟用户当下想做的事情高度相关。所以所有知识都在 skill 里,但 agent 还会带上那些文章作为一种短期记忆。我想跟所有人说,迭代你的 skill 是我个人做过的投入产出比最高的活动之一。所以你可以把 Pavel 的 skill 当成一个基线,这样你至少有个底子。然后当你遇到一些关于他这个 skill 的反馈时,把反馈给 Claude,说「我想让你改进我那个『识别现有假设』的 skill。读一下我们的聊天记录,看看我给你的反馈,搞清楚是什么根本原因导致你给出了那个让我不得不提意见的输出,然后从第一性原理出发重写这个 skill,让它不再犯那个错。」我发现这是我做的投入产出比最高的单项活动,所以我觉得一旦你拿到了初版 skill,你得真正去迭代它,才能榨出它的 ROI。我同意你,Aakash。我从不背 prompt,这就像在做 eval(评估)一样——你也写过关于 eval 的文章,说在搭建某个 AI 系统、某条 AI 流水线时,你得看系统在真实场景里表现如何,然后找出它的失败模式。在这个场景里你不用真去建一套 eval,你只要给 Claude 反馈、告诉它哪里错了就行。它能理解上下文,因为它已经有了这个
[35:55] Aakash
context. And yeah, what what were your expectations? And it will fix that and you you test it again, test it again and eventually we will eliminate maybe 99% of the of the failures. So that's that's the the only way and you cannot just sit and yeah, use some magic magic technique to to get it right on the first try. It doesn't work like that. So Pavel, now that people understand about self-improving skills, can you show us what our skill developed? So, we already the presented. Uh I just asked to design product strategy and it loaded this built-in So, it loaded two skills. One skill was designing presentations and we should see it. Uh yeah. PPTX. So, this is a skill by Entropic and another one was my skill about product strategy from my plugin. And it used it to to create the slide deck. We can see the directly in Coda. Um okay. I can see it in Google Drive or display it here. Directly here. So, we have product strategy canvas for Amazon. We have a product vision. Market segments like conscious customers, independent brands, um relative costs. So, Uh it adjusted colors. It figured out what the layout should be. This is not my branding. It's like uh Coda's invention. Also, icons and yeah, the the difference between orange and green. This is interesting. So, buyers have this whatever it means.
上下文。然后,对,「你的预期是什么?」它就会修好,然后你再测一遍、再测一遍,最终我们大概能消除 99% 的失败。这就是唯一的办法,你没法干坐着、靠什么神奇技巧一次就搞对,根本不是那么回事。那 Pavel,现在大家都明白了自我改进的 skill 是怎么回事,你能不能给我们看看我们刚开发出来的那个 skill 成果?我们其实已经演示过了。我刚才只是让它「设计产品战略」,它就加载了这个内置的——它加载了两个 skill。一个是「设计演示文稿」,我们应该能看到,对,PPTX,这是 Anthropic 出的一个 skill;另一个是我那个关于产品战略的 skill,来自我的 plugin。它用这两个来生成幻灯片。我们可以直接在 Cowork 里看。嗯,好。我可以在 Google Drive 里看,也可以直接在这显示。直接在这。所以我们有 Amazon 的产品战略画布。我们有产品愿景,有市场细分,比如有环保意识的客户、独立品牌、嗯、相对成本。所以——嗯,它调整了配色,自己琢磨出了布局该怎么排。这不是我的品牌色,是 Claude 自己发挥的。还有图标,对,橙色和绿色的区分。这个挺有意思。所以「买家有这个」,不管它什么意思。
[37:54]
[laughter]
[笑声]
[37:56] Aakash
It looks It looks like something smart. Uh tradeoffs. So, what we are not going to do. So, we have this focus on our strategy and the things that we are we want to say no to. So, uh makes sense. Key metrics. So, how we are going to track that our strategy is working? Like NPS, order value, seller retention. Seems to make sense. Uh North Star, a star icon. God like it is suggested guardrail and metrics. So, uh, when we focus on on our nodes, North Star, how what should we monitor to make sure that our our other areas do not degrade? Uh, and carbon neutral delivery rate. Uh, this can have an effect on Yeah, I'm not sure. Uh, like public relations. Our growth strategy, why in different market segments and unit economics. Like, this is pretty advanced. And it's not a single layout. It's not just Those are not just tables. It's uh, yeah, diverse layouts, diverse icons, diverse uh, it makes sense. Uh, if I ask a graduate to do that, like I I'm not sure I would get that in in a few hours hours. Yeah, so there's two really mind-blowing insights for everybody here. One, Claude is way better at using PowerPoint than it was a month or two ago. And so you need to look at this output and see like it can use PowerPoint to create great presentations. There's no excuse to walk into a meeting with a bad presentation anymore. And the second is that it used the skill, specifically some of the things Pavel had defined in the skill around having a North Star metric, have guardrails, and it has implemented those. And so that's why if you have a good skill, you can now create things. This is why people say rip McKinsey, right? You basically get a McKinsey level output in a minute or two. Yeah, I I have defined 60 something skills, but you can define how you can find [laughter]
看起来还挺聪明的。嗯,「权衡取舍」,也就是我们不打算做哪些事。所以这里是我们战略的聚焦点,以及我们想要说「不」的那些事。嗯,有道理。「关键指标」,也就是我们怎么追踪自己的战略是否奏效?比如 NPS、客单价、卖家留存率,看着挺合理。嗯,「北极星指标」,一个星星图标,挺好——它还建议了护栏指标。嗯,就是当我们聚焦在核心、聚焦在北极星上时,我们该监控什么,来确保我们其他领域不会退化?嗯,还有「碳中和配送率」,嗯,这个可能会影响——对,我也说不准——比如公关。我们在不同市场细分里的增长战略和单位经济模型,这相当高阶了。而且它不是单一布局,不只是——这些不只是表格,而是,对,多样的布局、多样的图标、各种各样的东西,挺合理的。如果我让一个应届生来做这个,我不太确定我能在几个小时内拿到这种成果。对,所以这里给大家有两个特别炸裂的洞察。第一,Claude 用 PowerPoint 的能力比一两个月前强太多了。所以你得看看这个输出,意识到它真的能用 PowerPoint 做出很棒的演示文稿,现在你再没借口带着一份糟糕的 PPT 去开会了。第二,它用上了 skill,具体来说就是 Pavel 在 skill 里定义的那些东西——比如要有北极星指标、要有护栏指标——它都实现了。所以这就是为什么只要你有一份好的 skill,你现在就能造出这些东西。这就是为什么大家会说「McKinsey 可以退场了」,对吧?你基本上一两分钟就能得到一份麦肯锡级别的输出。对,我已经定义了 60 多个 skill,但你能定义——你能找到[笑声]
[40:03] Aakash
hundreds of skills or thousands of skills defined by others. Uh, I would only recommend you to to verify every skill that is not defined by Anthropic. Uh, and make sure that it really fits your specific scenario, a specific use case. But other than that, yeah, the knowledge is uh there are many free repositories. I can also share those links with Akash, so So, we'll include those in our newsletters when we talk about this podcast, so that you guys get all of those. Be sure to subscribe to both. That's the Co-work segment. So, you guys just got the wow moment. Co-work can generate amazing presentations for you. We showed you how Co-work can organize your file system for you. Now, talk to us about CloudCode, Pavel. Why does a PM even need CloudCode? At this point, they think, "Oh, Co-work is good enough." Co-work is is not adjusted to working with code bases. And as a product manager, you will be working with code bases a lot. And also, if you are building complex systems that involve multiple files, then um this view that that what we see in Co-work is is not adjusted to it. So, for example, if I want to find this this presentation I can click show in folder and it is somewhere. Probably in our Yeah, we've been working with these invoices folder, so it places it there. Uh but imagine I have like 100 files or a few hundred files, like different invoices, different contractors, um contracts, maybe some some article drafts, some marketing strategies, pictures, brand guide, and so on. So, those files will be organized in some hierarchy. And you cannot browse them from here. You can just drag and drop or select single files. Like just by by browsing your local folder or browsing Google Drive. Um but there is no way to uh uh like easily work with with those large
几百个甚至几千个由别人定义的 skill。嗯,我只是建议你,凡是不是 Anthropic 定义的 skill,每一个都要核实一遍,确保它真的契合你具体的场景、具体的用例。除此之外,对,这些知识——有很多免费的 repo。我也可以把这些链接分享给 Aakash,所以——所以我们聊到这期播客时会把这些放进我们各自的 newsletter 里,让大家拿到所有这些资源。记得两边都订阅一下。这就是 Cowork 这一段。所以你们刚才已经看到了那个「哇」的时刻——Cowork 能帮你生成超棒的演示文稿,我们也给你们看了 Cowork 怎么帮你整理文件系统。那现在跟我们讲讲 Claude Code 吧,Pavel。为什么 PM 还需要用 Claude Code?这时候他们可能会想「噢,Cowork 已经够用了」。Cowork 并不适合处理代码库,而作为产品经理,你会大量跟代码库打交道。另外,如果你在搭建涉及多个文件的复杂系统,那 Cowork 里这种视图就不适合了。比如说,如果我想找到这份演示文稿,我可以点「在文件夹中显示」,它在某个地方,可能在我们的——对,我们一直在用这个 invoices(发票)文件夹,所以它就放那儿了。但想象一下我有上百个文件,几百个文件,各种不同的发票、不同的承包商、合同,可能还有一些文章草稿、营销策略、图片、品牌指南等等。这些文件会按某种层级结构来组织,而你在这里没法浏览它们,只能拖拽或者选中单个文件,靠浏览你的本地文件夹或者 Google Drive。嗯,但没办法轻松地处理那些大型的
[42:19] Aakash
codebases. One example, but also a list of contractors, a list of invoices, like a lot of graphical files, some templates. Uh you need to upload them, find those files, upload them here, and once the agent delivers something, you need to find that in the folder or or save it to the folder. So, even though this is my real folder, I still need to find this file uh to do something with it. Uh so, especially in case of uh codebases, uh you need this this view in which you see folders. You can expand folders, and you can see what is inside. So, for example, I have infographics folder, and inside infographics folder I have um Cloud Code Pricing, and inside Cloud Code Pricing I have some files generated by Cloud. So, let's see it in um I will just view it here. So, this is not the best picture that it generated. Uh yeah, but this was published today in my newsletter and it's also generated by Cloud. Uh another one can be this picture with calendar that went viral. This was also uh generated by Cloud, not by me. So, this one. This one is so good. So, a lot of people they struggle with getting faces and images and logos. So, you've gotten the Anthropic logo and you've gotten those eight people's faces. Also, you've actually gotten the data from Twitter. So, can you show us all three components of how you did that? Um yeah, I feel I think I need to explain how my system works and how it is organized to do that. Otherwise, it would be I can repeat the process, but but to explain how it works, I need to just to explain the system. Yeah.
代码库。这只是一个例子,但还有一份承包商清单、一份发票清单、一大堆图形文件、一些模板。嗯,你得上传它们、找到那些文件、传到这里,等 agent 交付了东西,你又得在文件夹里找到它,或者把它存回文件夹。所以即便这是我真实的文件夹,我还是得专门去找到这个文件才能对它做点什么。嗯,所以尤其是处理代码库时,你需要这种能看到文件夹的视图,你可以展开文件夹,看到里面有什么。比如说,我有个 infographics(信息图)文件夹,里面有个 Cloud Code Pricing(Claude Code 定价)文件夹,里面又有一些 Claude 生成的文件。我们来看看——我就在这直接看。这张不是它生成的最好的图,嗯,对,但这张今天发在了我的 newsletter 里,也是 Claude 生成的。嗯,另一张可以是这张带日历的图,那张当时挺火的,这也是 Claude 生成的,不是我做的。就这张。这张太棒了。很多人都搞不定人脸、图像和 logo,而你这张里搞定了 Anthropic 的 logo,搞定了这八个人的脸,你甚至还把数据从 Twitter 上扒了下来。那你能不能给我们演示一下这三个部分你都是怎么做到的?嗯,对,我觉得我得先解释一下我的系统是怎么运作、怎么组织的才行,不然就——我可以把过程复现一遍,但要解释它的原理,我就得把整个系统讲清楚。对。
[44:33]
[sighs]
[叹气]
[44:33] Aakash
Um So, recently Karpathy presented this um system in which you use LLMs to build a personal wiki or knowledge base for humans. So, you upload some information, give them give random articles and random attachments to to an agent documents, and agent organize them, and then you can browse those files, browse those information, and how different uh facts are connected. Uh I've been doing it since February 2026, and instead of building the second brain for myself, I started building second brain for my agents. Uh so so uh I'm the curator of the information. Um and what I do is I send articles, I send infographics that I find in on social media. I can also ask hey, let's analyze the last 10 posts by Akash um above 200 reactions. What why they worked. Voice um hooks um emotions And a lot of people struggle with editing their terminal prompts. You're using the to do that, right?
嗯,最近 Karpathy 展示了这么一套系统:你用 LLM 来给人类构建一个个人 wiki 或者知识库。你上传一些信息,把一些随机的文章、随机的附件、文档丢给一个 agent,agent 帮你把它们组织起来,然后你就能浏览这些文件、浏览这些信息,看不同的事实之间是怎么关联的。我从 2026 年 2 月就开始做这件事了,但我不是给自己建第二大脑,而是开始给我的 agent 建第二大脑。嗯,我是信息的策展人。我做的事情是:我把文章发进去,把我在社交媒体上看到的信息图发进去。我也可以说,「嘿,咱们分析一下 Aakash 最近 10 条、反响超过 200 个互动的帖子,看看它们为什么有效。」语气、嗯、钩子、嗯、情绪——还有很多人都搞不定怎么编辑他们终端里的 prompt,你是用——来做这个的,对吧?
[46:20] Aakash
Yeah, yeah. Yeah, there is the second way to use Cloud Code. I'm using CLI inside Visual Studio Code, but you can also use Visual Studio extension. Uh they are both similar and they are not super user-friendly.
对对。对,使用 Claude Code 有第二种方式。我是在 Visual Studio Code 里用 CLI,但你也可以用 Visual Studio 的扩展。两者都差不多,而且都不算特别好上手。
[46:36]
[snorts]
[嗤笑]
[46:36]
Uh okay, so yeah, in February I started um making screenshots on social media and I just I was giving different types of information to my agents. So, back then it was Co-work. And I asked what made this post work or what made this infographic work. And then the agent Co-work replied with with some information. In some cases it knew the answer, so I was able to to note it. In In other cases it replied with a hypothesis. So, this worked probably because something. Uh I decided that instead of noting this myself, I will ask an agent, "Hey, build a knowledge database and can you please every time I give you some article or infographic organize it by domain?" So, in this case the domain is social media, so this can be X, it can be LinkedIn, it can be Substack. And then uh write the rules for which you have a lot of information. Um if you see the repeating patterns, save this pattern as a rule and if you are not sure, save it as a hypothesis that you can later prove by analyzing more data. So, what we have ended up with is this knowledge database and I have um Yeah, for example, sound bites. So, what creators use to uh yeah, core patterns uh that work across platforms based on data. Uh what are the core techniques like interpretation layer? I'm comfortable clothes. Credibility before climate and this is confirmed across multiple many platform or platforms and all creators and all successful posts. How to What is the word choice? Like uh uh list of hypotheses across platform specific to to index rejected. So, also things that we have been considering in the past, but now we know they are not correct. Like let's say on cross platform, this is um Yeah. Achievement as a as proof hooks outperform achievement as point hooks or emotional diversification
呃好,是这样,二月份的时候我开始在社交媒体上做截图,然后我把各种不同类型的信息喂给我的 agent。那会儿用的还是 Cowork。我会问它:是什么让这条帖子火了?或者是什么让这张信息图奏效的?然后 Cowork 这个 agent 就会回我一些信息。有些情况下它知道答案,我就能记下来;另一些情况下它会给我一个假设,比如「这条可能是因为某某原因才有效」。后来我决定,与其自己手动记录,不如让 agent 来做。我跟它说:「嘿,建一个知识库吧,以后每次我给你一篇文章或者一张信息图,你能不能按领域帮我归类?」这里的领域就是社交媒体,可以细分成 X、LinkedIn、Substack。然后再去写规则——你已经有大量信息了,如果你看到重复出现的模式,就把这个模式存成一条规则;如果你不确定,就先存成假设,以后分析更多数据时再去验证。最后我们就得到了这样一个知识库。比如说里面有「金句」(sound bites)这一类——就是创作者们常用的、基于数据在各平台都通用的核心模式。还有核心技巧,比如「解读层」(interpretation layer)。「先建立可信度再谈观点」这一条是在多个平台、所有创作者、所有成功帖子上都得到验证的。还有用词选择,比如一份跨平台的假设清单,以及那些被否决的、特定于某个指标的条目。也就是说,有些我们过去考虑过、但现在知道是错的东西也会被记下来。比如在跨平台层面,「把成就作为佐证的钩子」优于「把成就作为重点的钩子」,或者「情绪多样化」
[49:14]
correlates with higher average engagement. Those are not hypotheses that I formulated. I have not even seen them. I'm seeing it for the first time this one. It's the numbers 46 and it's like looking at it on the file. Uh so, it it extracts this information and uh yeah, every time it analyzes more posts, more graphics, it tries to confirm or reject those hypotheses. And uh yeah, at the end of the day uh I have information how to write uh the files beliefs. For X, what are the hooks that work on X, what are that we monitor, what are the rules. Hard rules like um probably more technical ones. Uh and some templates that we know we have tested and we know that this template will work. Uh this doesn't mean that it writes for me. Um I previously demonstrated in the newsletter that like conversation that uh I suggest ideas, I suggest mistakes. So like raw knowledge and raw opinion, what I think about the specific news. But then Cloud adjust the format, adjust the style, adjust the hook um to the information, to the platform, and and what I'm trying to say to make it resonate with others. Uh about your Okay, about your posts. I'm not sure why Mhm. So is the X API free that you've hooked into? Or how did you Are you paying for it? I've not I haven't seen somebody connect into it before. Yeah, I'm paying for X API. So in in many cases you can fetch posts for free. Uh like it's this FX Twitter, something like that. I'm not sure why we are seeing errors. Usually it is not the case. So this um Yeah, but I saw API FX Twitter. Yeah. So this one is free, but uh it it is limited. It It doesn't always work. So just to save save costs, the agent tries to use it by default. Uh if it doesn't work, it uses this custom tool that we developed together. Like I just gave it
与更高的平均互动量相关。这些都不是我自己总结出来的假设,我之前甚至都没见过,这一条我现在还是第一次看到。它的编号是 46,我就这么在文件里看着它。所以它会把这些信息提取出来,然后每分析一批新的帖子、新的图,它就尝试去确认或否决这些假设。到头来,我手上就有了一套关于怎么写这些「文件信念」的信息:在 X 上,哪些钩子有效、我们在监控哪些东西、有哪些规则。还有硬性规则,可能更偏技术性那种。再加上一些我们已经测试过、确定有效的模板。但这并不意味着它替我写。我之前在 newsletter 里演示过,那种对话式的——我提供想法、我指出别人的错误,也就是说我给的是原始知识和原始观点,是我对某条具体新闻的看法。然后 Claude 再去调整格式、调整风格、调整钩子,让它适配信息本身、适配平台,以及我想表达的内容,让它能跟别人产生共鸣。呃说到你的——好吧,说到你的帖子,我不太确定为什么……嗯。你接进去的那个 X API 是免费的吗?还是说你怎么搞的,你是付费的吗?我之前没见过有人接进去过。对,我是付费用 X API 的。很多情况下你其实可以免费抓取帖子,就那个 FX Twitter 之类的。我不太确定为什么我们这会儿看到报错,平常一般不会这样。所以这个……对,但我看到了 API、FX Twitter。对,这个是免费的,但它有限制,不是每次都好用。所以为了省成本,agent 默认会先用它,如果不行,再用我们一起开发的那个自定义工具。我就是给了它
[51:41]
documentation of the Twitter API, and it created the tool for itself. So uh uh Uh it wraps this this API to something that is easy to use. Nice. Uh Uh okay, I'm not sure what what what One formula in Let me see what it talked about you. Pure analyst voice, radical topic diversity, he doesn't need to stay in PM, he can surf whatever wave is biggest on any given day because the mechanism revealed pattern transferred to food science, neuroscience, physics. Uh okay. That's the his top 10 has two tech posts and eight random curiosity posts. That's the opposite of Pavol Lane's strategy. The question isn't whether to copy some hypotheses, you already have it documented, the question is whether the mechanism revealed voice can work within your lenses. Uh I architecture agent design PM to link at the same scale. Like we have many of those conversations with Cloud Cotton coworker. Uh yeah, just to demonstrate. And similarly, I feed it with infographics and it tries to extract some patterns that catch the attention of all work in other creators. We select infographics that are easy to codify as HTML.
Twitter API 的文档,它就自己给自己造了个工具。它把这个 API 包装成了一个很好用的东西。不错。呃好,我不太确定那个……Let me see,让我看看它说了你什么。「纯粹的分析师口吻,话题极度多元,他不需要固守 PM 领域,哪天哪个浪头最大他就冲哪个,因为这种『揭示机制』的模式可以迁移到食品科学、神经科学、物理学。」呃好。「他的前十条帖子里只有两条是科技类,其余八条都是随机的好奇心话题。这跟 Pavol Lane 的策略正好相反。问题不在于要不要照搬某些假设——你已经把它记录下来了——问题在于这种『揭示机制』的口吻能不能在你自己的视角框架里跑通。」呃,我把 agent 设计、PM 都串到同一个尺度上。我们跟 Claude、Cowork 有很多这样的对话。对,就是演示一下。同样地,我也喂给它信息图,它会尝试提取那些能吸引注意力的、在其他创作者那儿也起作用的模式。我们专挑那些容易用 HTML 编码出来的信息图。
[53:15]
[laughter]
[笑声]
[53:15]
And when they are easy to codify, then we design we extract components that we can reuse. And then it uses the growing library of components to design infographics for Does it make sense? Yeah, it does. The last component is how do you personally Did it pass fetch the profile pictures of Boris and Tariq and Uh, uh, this one? Yeah, but it was like a custom custom query. Like to I'm not sure what it was. Uh, it was ad hoc ad hoc maybe article, maybe something else. Let me see. Uh, calendar? No, assets, no. Uh, maybe drafts. No. It was like a temporary artifact. They asked it to create a script and go through uh first uh find so I knew the free Anthropic accounts. So, then I asked it to analyze their recent posts and reposts and I assumed that they will eventually retweet every other uh cloud code team member. Then it analyzed which team members used we or we just released or my team just released uh your if it was unsure it verified the comments so that yeah, just to make sure that the person was was a Anthropic employee. And then yeah, it went through like 15 people from Anthropic uh mapping for every feature mapping who first wrote about it. Um, for for those people we got pictures from also from Twitter from X. And then we iterated several times on how to how to visualize it. So, that was the final result but this is just like done by Cloud Code. This this particularly this one was done by Cloud Code and this is a HTML uh based on the research and the the the most difficult part was the research. So, just getting all this information from Twitter. Yes, it was difficult, but in the end you were just you were talking to Claude code over hours in natural language. So, it's not impossible for other people to do. Yeah, I do not code at all.
等到它们容易编码出来之后,我们就设计、提取出可复用的组件。然后它就用这个不断增长的组件库去设计信息图。这样讲得通吗?嗯,讲得通。最后一个环节是你个人怎么——它有去抓 Boris 和 Tariq 他们的头像吗?呃这个?对,但那是个自定义查询,具体我不太确定是什么,是临时拼的,可能是篇文章,也可能是别的什么。让我看看。呃,日历?不是。Assets,不是。呃,可能在 drafts?不是。它是个临时 artifact。我让它写个脚本,先去找——我本来就知道那些免费的 Anthropic 账号,然后我让它分析他们最近的帖子和转发,我假设他们最终都会互相转发每一位 Claude Code 团队成员的内容。接着它分析哪些团队成员用了「我们」「我们刚发布了」或者「我团队刚发布了」这种说法,如果不确定,它就去核实评论区,确保这个人确实是 Anthropic 的员工。然后它就这么过了大概 15 个 Anthropic 的人,针对每个功能去映射出是谁最先写到它的。对这些人,我们也从 Twitter、从 X 上拿了头像。然后我们反复迭代了好几次,琢磨怎么把它可视化。这就是最终成果,但这完全是 Claude Code 做的。这一个特别是 Claude Code 做的,是基于研究生成的 HTML,最难的部分就是研究本身,也就是从 Twitter 上把所有这些信息扒下来。对,是挺难的,但说到底你只是在用自然语言跟 Claude Code 聊了好几个小时,所以这对其他人来说也不是做不到。对,我完全不写代码。
[55:53]
Like I I don't write any code. I I don't even review the code. Uh if I want to know how something works, I just ask questions in the chat window. So, how do you make this system self-improving and what are the tips and tricks that people need to know around Claude MD files and folder structure to make their system really sing? Okay, so first s- a lot of people uh so people are uh discussing cloud.md file that you can put your instructions there, but but the problem with cloud.md is that uh if you put all the instructions inside, it it will grow maybe not exponentially, but it will keep growing and growing and growing and eventually it will consume a lot of your of your context window and every time you ask a simple prompt in your project, all this cloud.md context will be included. Uh it is it is part of your prompt. Uh so, a smarter approach is to organize your knowledge in uh files dedicated to specific domains. Let me let me demonstrate mine. Uh so, I have cloud.md. So, I have this main cloud.md. So, the only goal of cloud.md is to explain what this project is about. It doesn't have detailed instructions or detailed information about uh like good practices or bad practices of what to avoid, what to do more often. Uh all this information is inside uh in other files and the only way a goal of Cloud MD is to to give those instructions how to find the knowledge and what to do with the new knowledge. Uh so, project structure, this is just what we can see on the left side. So, I can skip that. Like there are tools, there are scripts. Um just so that the agent knows without scanning the the repo. Uh another one uh is where things live. So, this is similar. Uh And by the way, this was also created by Cloud. I I didn't wrote this. So, I I don't discuss in the chat how we
我一行代码都不写,连代码都不审。如果我想知道某个东西是怎么运作的,我就在聊天窗口里直接问。那你怎么让这套系统能自我改进?围绕 CLAUDE.md 文件和文件夹结构,有哪些技巧是大家需要知道、才能让自己的系统真正跑顺的?好,首先——很多人都在讨论 CLAUDE.md 这个文件,说你可以把指令放进去。但 CLAUDE.md 的问题在于,如果你把所有指令都塞进去,它会不断变大——可能不是指数级增长,但会越长越大、越长越大,最终吃掉你大量的 context window。而且每次你在项目里发一个简单的 prompt,整个 CLAUDE.md 的内容都会被带上,它是你 prompt 的一部分。所以更聪明的做法是,把你的知识组织到一个个专门对应某个领域的文件里。我来给你演示一下我的。呃,我有 CLAUDE.md。我有这个主 CLAUDE.md,它唯一的目的就是说明这个项目是干什么的。它不包含详细指令,也不包含关于好做法、坏做法、该避免什么、该多做什么的详细信息。所有这些信息都在其他文件里,CLAUDE.md 的唯一目标就是告诉它怎么去找到这些知识,以及拿到新知识后该怎么处理。呃,项目结构这部分,就是你在左边能看到的那些,我可以跳过。里面有工具、有脚本,就是让 agent 不用扫描整个 repo 就知道有这些东西。呃,还有一个是「东西放在哪里」(where things live),跟前面类似。顺便说一句,这也是 Claude 写的,不是我写的。我不会在聊天里讨论我们怎么
[58:08]
organize things and Cloud writes instructions for itself. Uh who I am, like basic information. Uh and uh the knowledge system, this is the most important part. So, this is the self-organizing system where uh I have index of all files, like X files, LinkedIn files, Substack files, Substack notes files. Uh voice archetypes, craft. So, those are This is the general knowledge across platforms. And then, for every platform, where different uh elements live. And now, uh there are also workflows that the system should follow. So, how to fetch data for Twitter, how to how to uh work with LinkedIn posts, how to link with work with Substack. Uh The most important part is when asked to study and analyze. So, every time I give it some posts, like let's study the last 100 posts or I like this one, uh let's analyze what made it work. Uh it knows what tools to use. So, it will extract hook pattern structure, sound sound bites, and engagement metrics. Uh, only if I request visual analysis, it will also analyze the attached graphic. This is to save tokens. So, I I don't analyze the graphics every time I analyze the post. Uh, check against the existing patterns and false beliefs. And then, if if there is an existing hypothesis, it will update the existing hypothesis with the new evidence. If there is an existing hypothesis, but we see that the specific post didn't work, we can demote the hypothesis that the hypothesis is less likely or it will it will become rejected. Uh, yeah, and it will append the post that was analyzed to the to the database. So, uh, I think we may need to simplify this. Uh, but basically, it gets information about posts, and it tries to decompose this knowledge, and organize this knowledge itself into rules, into hypotheses, and yeah, like other elements like
组织东西,是 Claude 自己给自己写指令。呃,「我是谁」,就是些基本信息。还有知识系统,这是最重要的部分。这是一套自组织系统,里面我有所有文件的索引,比如 X 的文件、LinkedIn 的文件、Substack 的文件、Substack notes 的文件。呃,还有「声音原型」(voice archetypes)、「手艺」(craft),这些是跨平台的通用知识。然后针对每个平台,又分别列出不同的元素放在哪里。还有一些系统应该遵循的工作流,比如怎么为 Twitter 抓数据、怎么处理 LinkedIn 帖子、怎么处理 Substack。最重要的部分是「当被要求研究和分析时」。每次我给它一些帖子,比如「研究一下最近 100 条帖子」,或者「我喜欢这条,分析一下它为什么有效」,它就知道该用哪些工具。它会提取钩子、模式、结构、金句和互动指标。只有当我明确要求做视觉分析时,它才会去分析附带的图片,这是为了省 token。所以我不是每次分析帖子都去分析图。呃,然后它会对照已有的模式和错误信念做检查。如果有现成的假设,它就用新证据去更新这个假设;如果有现成的假设、但发现这条具体的帖子没奏效,它就给这个假设降权——让这个假设变得不那么可能,或者直接变成被否决。呃,对,然后它会把分析过的帖子追加到数据库里。所以呃,我觉得这套可能需要简化一下。但基本上,它拿到关于帖子的信息,就尝试把这些知识拆解开,自己把知识组织成规则、组织成假设,还有其他元素,比如
[1:00:35]
structure, hooks, sound bites, uh, engagement metrics, and so on. So, the system learns itself without me telling it, uh, why the specific piece of content worked. Uh, the cloud cloud does it itself. And the next time I have an idea or I express opinion about something that the other person tweeted, then I can say, "I like this, but I think this this is connected to another idea." Or, for example, I like Karpathy post, but uh, it will degrade over time. And by the way, uh, you don't need to obsidian where if the person if the user is not a person, because we are building a knowledge base for agents. And yeah, how can we retweet that or maybe I can add some additional details and then Cloud it will suggest it will use my ideas it will it will use my my tag but it will format it in a way that resonates. So this is very content focused. Yeah. A PM watching this might say this is kind of content focused. What's in this for me and I don't write code. So what is the minimum viable setup for a product manager and how they should be setting this up? Uh self-improving system? Um that's uh yeah, I have recently shared that and let me demonstrate it. So I have actually I have a poster for this. So how to create a knowledge system that learns itself and the prompt is very simple. So you don't have to build everything I did. You can start with just this prompt.
结构、钩子、金句、互动指标等等。所以这套系统是自己在学习,不需要我去告诉它某条具体内容为什么有效,Claude 自己就把这事做了。下一次我有了想法,或者我对别人发的某条推表达看法时,我就可以说「我喜欢这条,但我觉得它跟另一个想法是有联系的」,或者比如「我喜欢 Karpathy 这条帖子,但它会随时间贬值」。顺便说,你不需要 Obsidian——如果使用者不是真人的话,因为我们是在为 agent 构建知识库。然后,我们怎么转发那条、或者我加点额外细节进去,接着 Claude 就会提建议、会用我的想法、用我的标签,但它会用一种能引起共鸣的方式重新排版。所以这套很偏内容。对。一个看这期节目的 PM 可能会说:这太偏内容了,这对我有什么用?我又不写代码。那对一个产品经理来说,最小可行的搭建方案是什么?他们该怎么把这套搭起来?呃,自我改进的系统?嗯,这个——对,我最近分享过,我来演示一下。我其实专门为这个做了一张海报,讲「怎么创建一个能自我学习的知识系统」,而且这个 prompt 非常简单。你不需要把我做的全套都搭出来,从这个 prompt 开始就行。
[1:02:17]
[laughter]
[笑声]
[1:02:17]
So before starting a new task uh review um existing rules and hypothesis for this domain. Then apply rules by default. So for example, if if we know that uh So basically this is the most important part that you need to paste to your Cloud and D. that uh and this is not content specific. So that whatever Cloud does something in in a specific domain like testing software like uh writing marketing materials, maybe writing release notes, it should uh review the rules and hypothesis from this domain and it should apply the rules that were confirmed to its work. So for example, how a good test case, how good user stories formatted or how how the acceptance criteria should be written or uh what are the good examples of of release notes or customer offer whatever you do. What are the good examples of something and what are the bad examples of something and it it is it will try to extract the rules. And then when you ask it to perform a task like hey you saw the 10 good examples to to bad examples then let's try to to create a another offer for for this new customer. It will review the existing rules and hypotheses because you didn't write it. It reasoned what are the good rules and bad rules, what are the hypotheses from the data that it it has seen and it will also keep learning. So every time you give it a new information, it will update its knowledge, it will update the hypotheses. It can also ask you questions if this is something that where you can give agent a feedback. Uh Then basically it keeps learning, keeps keeps adding knowledge and the knowledge is organized by the domain so pricing, so marketing, testing, quality, strategy. Yeah, so you want to create that knowledge with an index.md that has a router and in your cod.md you want to give it this prompt
就是说,在开始一个新任务之前,呃,先回顾一下这个领域已有的规则和假设。然后默认应用这些规则。比如说,如果我们知道——基本上这就是你需要粘到 CLAUDE.md 里最重要的那部分,而且它跟具体内容无关。这样无论 Claude 在某个具体领域做事——比如测试软件、写营销材料、可能写发布说明(release notes)——它都应该回顾这个领域的规则和假设,并把那些已被确认有效的规则应用到工作中。比如说,一个好的测试用例长什么样、好的用户故事怎么写、验收标准该怎么写,或者发布说明、客户报价的好例子是什么——不管你做的是什么。某样东西的好例子是什么、坏例子是什么,它就会去尝试提取出规则。然后当你让它执行一个任务时,比如「嘿,你看过那 10 个好例子和坏例子了,那咱们试着给这个新客户再做一份报价」,它就会回顾已有的规则和假设——因为这些不是你写的,是它根据自己看过的数据推理出来:哪些是好规则、哪些是坏规则、有哪些假设——而且它还会持续学习。每次你给它新信息,它都会更新自己的知识、更新假设。如果有什么地方可以由人给 agent 反馈,它也会反过来问你问题。呃,然后它基本上就一直学、一直加知识,而且知识是按领域组织的,比如定价、营销、测试、质量、策略。对,所以你要做的就是创建这样一套知识,配上一个带路由的 index.md,然后在你的 CLAUDE.md 里给它这个 prompt,
[1:04:35]
so that it's self-improving. It will figure out what to do and every time it it's every time you you ask it to do something it will use the existing knowledge that it is growing and every time it it sees new information with the context like when I quote a tweet, the tweet has some metrics like this tweet worked or didn't work or how many people liked it. Uh but uh maybe you can feed it with with the offers that worked or with Um yeah. Resumes of successful candidates and it will start generating those rules and the next time you will get a candidate you can ask Cloud hey is this a good candidate? Love it. So [snorts] one of the things that you've written about that I haven't seen a lot of people write about and maybe you can show us is the Chrome MCP. When and how should we be using that? I don't use Chrome MCP anymore. And the reason is that I've been testing different approaches and Chrome MCP it's basically MCP that controls your browser. It works well probably similarly to Cloud in Chrome which is Anthropic extension that you can use like this guy here. You can ask to do something or you can also call it from you can schedule tasks here or you can call to to co-work and co-work will will call this Cloud in Chrome. The problem with those extensions is that they rely heavily on taking screenshots. And screenshots mean a lot of tokens and if you have some tasks that you want to repeat regularly or complex processes the yeah you can easily consume like $100 in an hour especially with with Opus. So what I do instead is I use let me check but I think my agents right now use agents dev. So what agent browser does by Vercel Labs this is the most reliable according like in my tests. It also uses the real browser but it can do that in a headless mode and it explains
这样它就能自我改进。它会自己想清楚该做什么,每次你让它做事,它都会用上那套不断增长的现有知识;每次它看到带 context 的新信息——比如我引用一条推,这条推带着一些指标,说明它火了或者没火、有多少人点了赞——呃,或者你也可以喂给它那些成交过的报价,或者,嗯,那些成功候选人的简历,它就会开始生成那些规则,下次再来一个候选人,你就可以问 Claude「嘿,这是个好候选人吗?」太喜欢了。所以呢,你写过一个东西,是我没怎么见别人写过的,也许你可以给我们演示一下,就是 Chrome MCP。我们应该在什么时候、怎么用它?我现在已经不用 Chrome MCP 了。原因是我一直在测试不同的方案,而 Chrome MCP 本质上就是一个控制你浏览器的 MCP。它运作得不错,可能跟 Claude in Chrome 差不多——那是 Anthropic 的一个扩展,你可以像这位老兄这样用,你可以让它做点什么,也可以从这儿调它,你可以在这儿安排定时任务,或者调用 Cowork,让 Cowork 去调 Claude in Chrome。这些扩展的问题在于,它们高度依赖截图,而截图意味着大量 token。如果你有些任务要定期重复,或者是复杂流程,那你很容易一小时就烧掉 100 美元,尤其是用 Opus 的时候。所以我现在改用——让我确认一下,但我觉得我的 agent 现在用的是 agent 的那个……(agents dev)。所以 Vercel Labs 做的这个 agent browser,根据我的测试是最可靠的。它同样用真实浏览器,但能在无头(headless)模式下运行,而且它会解释
[1:07:08]
the structure of the page to the agent without presenting the entire HTML. So, this is it is token efficient. And so, the agent doesn't have to see HTML and it doesn't have to interpret HTML, but can take actions. So, it will see buttons with specific IDs even if the original button didn't have ID. So, this is a real browser, real rendering. It can execute JavaScript if needed. Uh it waits for rendering, which is important. So, if this is a single page application or there are some additional resources that load a few seconds later, it will wait for it. And yeah. It presents this page to an agent without taking screenshots. So, the agent can see the text, it can see layout, it can see different components, the content. It can say, "Hey agent browser, click this button." But, it doesn't have to interact with parse or interact with HTML itself. It is very simple simple protocol. It's not an MCP, it's a CLI tool. Very cool. Use agent browser from Vercel instead of Chrome MCP. And what do you use it for exactly? Like, what is the use case that a PM should think about? This is when I
把页面的结构解释给 agent,而不用把整个 HTML 都呈现出来。所以它很省 token。这样 agent 既不用看 HTML、也不用解读 HTML,但照样能采取行动。它能看到带特定 ID 的按钮,哪怕原始按钮本来没有 ID。所以这是真实浏览器、真实渲染,需要时还能执行 JavaScript。呃,它会等待渲染完成,这点很重要。如果这是个单页应用,或者有些额外资源是几秒后才加载的,它会等它们加载好。对。它把这个页面呈现给 agent,但不用截图。所以 agent 能看到文字、能看到布局、能看到不同的组件和内容。它可以说「嘿 agent browser,点这个按钮」,但它不用去解析或直接操作 HTML。这是一套非常简单的协议,它不是 MCP,是个 CLI 工具。很酷。用 Vercel 的 agent browser,别用 Chrome MCP。那你具体拿它来干嘛?一个 PM 应该想到的使用场景是什么样的?这个嘛,当我
[1:08:34]
Like, everything where you do not have an API to get data from external systems. So, for example, if I want to get data from LinkedIn, of course, I can fire Chrome MCP or trigger cloud in Chrome, but it will start taking screenshots every half a second or every second. So, it's like crazy token consumption. It it also visible in most cases on your screen, so it's not something that uh you would like to see. And, with this I can just ask it to hey go to LinkedIn, check my uh inbox, or analyze the top last posts by Akash. And, see the comments, so yeah, something like that. So, maybe more enterprise use case will be accessing some legacy software without API or without MCP uh servers. Uh maybe SAP. Maybe some old CRM system. Uh Yeah, yeah, but the data can access that just by using uh web browser. This is pretty epic, so as a PM, you need to hook your cloud code into absolutely everything, and you should use agent browser for the stuff where you don't have an MCP, CLI, API to hook into. Now, I want to move into remote work. Anthropic shipped four remote surfaces in recent months. Web sessions, remote control, dispatch, and channels. You use all four of these remote methods. Walk us through how you actually combine them for the same project, and when a PM should use which. Okay. [laughter] So, the um one issue with Anthropic is that those surfaces overlap, and I I don't use them in the same proportions. Uh so, I have tested channels. Uh I have abandoned channels because I I didn't see value like having this a Telegram interface because I have cloud um I have a dedicated cloud sub where I can do everything. Okay, but starting with maybe let's start with dispatch. So, uh dispatch is uh, a new tab that appeared in the desktop app and also on
基本上就是所有那些你没有 API 能从外部系统取数据的场景。比如说,我想从 LinkedIn 拿数据,当然我可以启动 Chrome MCP 或者触发 Claude in Chrome,但它会每半秒或每秒就截一次图,token 消耗简直疯了。而且大多数情况下你屏幕上还能看到它在动,这不是你愿意看到的。而用这个,我就可以直接让它「嘿,去 LinkedIn,看看我的收件箱」,或者「分析一下 Aakash 最近的几条帖子,看看评论」,类似这样。所以,更偏企业的使用场景可能是访问那些没有 API、没有 MCP 服务器的遗留软件,呃,比如 SAP,比如某些老旧的 CRM 系统。呃,对对,但这些数据光靠用 web 浏览器就能访问到。这太牛了。所以作为 PM,你得把你的 Claude Code 接到几乎所有东西上,而对于那些你没有 MCP、CLI、API 可接的部分,就用 agent browser。现在我想转到远程工作这块。Anthropic 最近几个月发布了四种远程入口:web sessions、remote control、Dispatch 和 channels。这四种远程方式你都在用。给我们讲讲你实际上是怎么把它们组合到同一个项目上的,以及一个 PM 该在什么时候用哪一个。好。[笑声] 所以呢,Anthropic 有个问题,就是这些入口彼此重叠,而我用它们的比例也不一样。呃,channels 我测试过,但我已经弃用 channels 了,因为我没看出有什么价值——搞这么个 Telegram 式的界面,因为我已经有 Claude……我有一个专门的 Claude 频道(sub),在那儿什么都能干。好,那也许我们先从——咱们先从 Dispatch 说起吧。呃,Dispatch 是桌面端 app 里新出现的一个标签页,在……上也有
[1:11:00] Pavel
your phone and this is exactly the same. So, here um, uh, okay. I'm not sure what it is. Uh, So, this part is like this single single interface in which you can interact with uh, cloud called and co-work. It is displayed under under co-work. It It is a bit confusing. It is a completely different product. It's like walkie-talkie. So, uh, it can start multiple background tasks and every time the task is completed, it will report back what is the status. So, for example, uh, create create an infographic in Anthropic style for the following text. And let me find something. Um, following text and the text will be about board and windsurf and some some it's probably your work, so uh, in Anthropic style for the following text. And this will be uh, LinkedIn resolution. Uh, then I can ask it to hey, how many emails um, in the last 2 hours did I receive? No names, no personal data. So, like and it doesn't have to start the previous task, it will just start another one, and if we open the recents, I will show them. We should see that uh it delegates them to It should delegate them to to other threads. Uh three emails in the last couple of hours, so maybe it did it directly, and let's go to code. Ah. Analyze Akash posts. No normally it uh it it would delegate those tasks to subagents, and you will see them as dispatched tasks in the left panel. So, you can start multiple agents and use a single interface to communicate with all of them. Uh and I can take my phone, open the Let me open the cloud app. And I will try to demonstrate that. Uh but first, I will ask it Are you there? I'm not sure this will be visible, but like I have the same interface here. Yeah, I can see it. And so, you can basically talk to it on [clears throat] web and mobile now, and
在手机上是一模一样的。所以,这里嗯……呃,好吧,我不太确定这是什么。呃,所以这一部分就像是一个单一界面,你可以在里面跟一个叫 Cowork 的东西交互。它显示在 Cowork 下面,是有点让人困惑——它其实是一个完全不同的产品,更像是对讲机。所以呃,它可以启动多个后台任务,每次任务完成时它都会回报当前的状态。比如说,呃,「用 Anthropic 风格为下面这段文字做一张信息图」。让我找点东西。嗯,下面这段文字会是关于 board 和 windsurf 的,还有一些……大概是你的工作内容,所以呃,用 Anthropic 风格为这段文字做。这会是 LinkedIn 的分辨率。呃,然后我可以问它,嘿,「过去 2 小时里我收到了多少封邮件?不要名字,不要个人数据。」所以呢,它不需要先结束上一个任务,它会直接再启动一个,如果我们打开 recents(最近列表),我会把它们展示出来。我们应该能看到它把这些任务委派给了……它应该会把它们委派给其他线程。呃,过去几小时里有三封邮件,所以可能它是直接处理掉了,我们去 code 看看。啊,「分析 Aakash 的帖子」。不,正常情况下它会把这些任务委派给 subagent,你会在左侧面板里看到它们作为 dispatch 的任务出现。所以你可以同时启动多个 agent,并用一个界面跟它们全部沟通。呃,我可以拿起我的手机,打开……让我打开 Claude app。我会试着演示一下。呃但首先,我会问它「你在吗?」我不确定这能不能看清,但就像,我这里有同样的界面。对,我能看到。所以你现在基本上可以在网页和手机上跟它对话了,而且
[1:14:23] Aakash
so there's really no excuse. You're out on a walk, you're at a meeting, whatever it might be, you're at a conference, you can have your agents running for you. Yeah. And what about code web sessions? You said you use those roughly 60% of the time in dispatch 40%? Yeah, I don't remember the exact the exact stats. Like I really use all the all three surfaces, uh and the proportions differ from day to day. Um but most of the time I use dispatch and web sessions. Um maybe like 70% together. Uh then 5% chat sometimes, and the rest is cloud code. And the reason why I use dispatch so much is that I just don't work with my laptop. I I go for a shopping, I go somewhere with my kid, and I just dispatch tasks. Uh I already explained that I don't code. I don't I just provide feedback text feedback in the chat. And then Co-work dispatch presents me the results. I look at the results, I dispatch another task, and then I can continue what what I was doing. So, yeah, it it really transformed my transformed my my ways my days. Amazing.
所以真的没什么借口了。你在外面散步、在开会、不管在哪、在参加会议,你都可以让你的 agent 替你干活。对。那 Claude Code 的网页会话(web session)呢?你说过你大概 60% 的时间用那个,40% 用 dispatch?对,我不记得确切的……确切的数据了。我其实三个界面都用,呃,比例每天都不一样。嗯,但大部分时间我用 dispatch 和网页会话。嗯,加起来可能 70% 左右。呃,然后 5% 偶尔用 chat,剩下的是 Claude Code。我之所以这么频繁用 dispatch,是因为我根本不带着笔记本电脑工作。我去逛街、带孩子出门,我就直接 dispatch 任务。呃我前面已经说过我不写代码,我只是在 chat 里给反馈,文字反馈。然后 Cowork dispatch 把结果呈现给我。我看结果,再 dispatch 另一个任务,然后我可以继续做我手头的事。所以,对,这真的改变了我的……改变了我做事的方式,改变了我的每一天。太棒了。
[1:15:43] Pavel
Yeah, we have not covered web sessions. So, web sessions are something different. Uh so, here in Co-work in dispatch depending on the task that you give it, it can either dispatch task to Co-work like I asked about last three tweets and you see there is a dispatch thread here. So, this is something where I do not have to interact with it. It's dispatch talking to this agent. But it is visible for humans. And similarly, I can also dispatch some coding tasks from dispatch and it will it will implement that and report back. Uh Oh, it even created a graphic for your post like like let's say AI prototyping graphic. Ooh, I want to see it. Uh open like Wow, pretty good. Yeah, sometimes it takes a few but it is not the the the worst point to start. Very cool. Uh Okay, so this is the dispatch. So you can think of dispatch. Normally you don't use it on on your desktop. You you have the single chat, single interface that you use on mobile. And but sometimes um but go for dispatch to work, your computer must be online. And sometimes this is not the case, and sometimes there is a problem, and it stopped stops working. Uh or sometimes you have so many parallel threads that it is difficult to manage them from a single chat interface. There are no tabs here. Uh on your phone, you will see just one big stream of messages. You cannot organize them. Uh so when I do a more complex work, I switch to this code task. And this is you can think of it as uh Visual Studio Code and Cloud Code. It's just in the cloud, so it is hosted by Entropic. Um so it's yeah, it's just just a little a list of sessions where I can select a specific folder. Either my local folder or folder in GitHub, like my editor project. It is synced with GitHub, so all those files, all knowledge files
对,我们还没讲网页会话呢。网页会话是另一回事。呃,在 Cowork 里的 dispatch 中,根据你给它的任务,它可以把任务 dispatch 给 Cowork——比如我刚才问的「最近三条推文」,你看这里就有一个 dispatch 线程。这是我不需要去交互的东西,是 dispatch 在跟这个 agent 对话,但人是能看到的。同样地,我也可以从 dispatch 里 dispatch 一些写代码的任务,它会去实现并回报。呃,哦,它甚至还给你的帖子做了一张图,比如说「AI 原型设计图」。哦,我想看看。呃,打开一下……哇,挺不错的。对,有时候要等一会儿,但作为一个起点也不算最糟。很酷。呃,好,这就是 dispatch。你可以把 dispatch 理解成——通常你不会在桌面上用它。你用的是那个单一的 chat、单一界面,在手机上用。但有时候……不过要让 dispatch 工作,你的电脑必须在线。而有时候并不是这样,有时候会出问题,它会停止工作。呃,或者有时候你有太多并行的线程,单一的 chat 界面很难管理。这里没有标签页。呃在手机上你看到的就是一大串消息流,你没法整理它们。呃所以当我做更复杂的工作时,我就切换到 code 任务。这个你可以理解成呃 Visual Studio Code 加上 Claude Code,只不过它在云端,由 Anthropic 托管。嗯,所以,对,它就是一个会话列表,我可以在里面选一个特定的文件夹——要么是我本地的文件夹,要么是 GitHub 里的文件夹,比如我的 editor 项目。它跟 GitHub 同步,所以所有那些文件、所有的知识文件
[1:18:20] Pavel
uh hypotheses, uh sound bites, hooks, this is all synced with my private GitHub repository. And and even when my laptop is offline, I can go here from my mobile phone and ask it some question, and it will it will work in the cloud with without any device. Yep, and this is actually a more secure way to run stuff, guys, cuz it's in an Entrotic servers. So, I highly highly highly recommend everything you build, all your operating systems, you put them into GitHub, you point them via code web sessions, and you work with stuff here. You can also The coolest thing is we just showed with Dispatch is you can start a chat here, start something here on your desktop, and then you can go to your phone. And so, it's like you can be working 24/7 with us. Yeah, that's correct. And I I I don't work 24/7, but like I feel that my life is now much more works much much better integrated with life, so I don't have to have these blocks dedicated to work. I can go on a shopping and yeah, maybe when uh when shopping the this person task and then check the minutes later when I have a graphic, provide some feedback, and then continue what what I was doing. Uh so, this is really transforming, and when you need a better organization, or when you know that you want to focus on coding, like my Accridia.io platform, mm then I switch to code, and I also do it primarily on mobile. Uh and yeah, like in some cases I I also use code work and and Visual Studio Code, but primarily this is remote via most of the work is remote work. This is epic, guys. So, hopefully you can understand how your life will change once you set all this up. I want to draw out some key lessons, Pavel, from all of the hundreds of hours you've now put into these things. Starting with what's the biggest mistake PMs make when
呃,假设、金句、钩子,这些全都跟我的私有 GitHub 仓库同步。而且即使我的笔记本电脑离线,我也可以从手机上来到这里问它一些问题,它会在云端工作,不需要任何设备。对,而且这其实是一种更安全的运行方式,各位,因为它在 Anthropic 的服务器上。所以我非常非常非常推荐,你构建的一切、你所有的操作系统,都放进 GitHub,通过网页会话指向它们,然后在这里干活。你还可以……最酷的一点是,我们刚才用 dispatch 展示了——你可以在这里开一个 chat,在桌面上启动一些东西,然后切到手机上继续。所以这就像,你可以全天候 24/7 跟我们一起工作。对,没错。我我我并不是真的 24/7 工作,但我感觉我的生活现在跟工作的结合好太多了,所以我不必再划出一块块专门工作的时间。我可以去逛街,然后呃逛街的时候 dispatch 这个任务,几分钟后等有了图再来看看,给点反馈,然后继续我手头的事。呃,所以这真的很有变革性。当你需要更好的组织时,或者当你知道你想专注写代码时——比如我的 Accredia.io 平台——嗯那我就切换到 code,而且我主要也是在手机上做。呃,对,有些情况下我也会用 Cowork 和 Visual Studio Code,但主要是这样……大部分工作都是远程的。这太赞了,各位。希望你们能理解,一旦你把这一切都搭好,你的生活会怎样改变。我想提炼一些关键经验,Pavel,从你为这些东西投入的这几百个小时里。先从这个开始:PM 在……时犯的最大错误是什么
[1:20:26] Pavel
setting up cloud? Yeah, I think that the biggest mistake would be to prompt it every time from scratch instead of using Cloud. I force myself to organize knowledge and this is very difficult to me when writing articles and but Cloud can do it without effort. So, instead of collecting prompts and figuring out how to do something it's much easier to just to build a system where Cloud can learn from its mistakes and either from from your feedback or from data and it can figure out hypothesize and figure out the better ways to to the work. Yeah, not doing that, not organizing your knowledge, not learning from from mistakes and just having everything in your head and yeah, this is the the biggest mistake PM can make. Just come hoping that you can learn the better prompts. You can there's too much data to analyze. Cowork versus Cloud Code. If a PM only had a little bit of time to learn one, which one should they choose and why? This is not an alternative. I would start with Cowork because everything you you will learn in Cowork will help you better understand code. I use the same repo from Cowork and from from Cloud Code. So, for example, I already presented that but once again, I can select this editor project. Maybe let's remove this one. Um and it has Cloud MD, it will see the same files. It will So, this this is Cloud MD that I presented in Visual Studio. It will see the same files, the same structure. So, I can ask it to analyze tweets here in this interface. I can do it on my mob in the using this patch. I can do it in web session web Cloud session. I can do it in Visual Studio. This is all to the same repo. Uh I will start with Cogram because this interface is more simple. You don't have to uh get used to uh explorer and terminal.
……在搭建 Claude 的时候?对,我觉得最大的错误就是每次都从零开始 prompt,而不是用 Claude。我逼着自己去组织知识,这对我来说在写文章时非常困难,但 Claude 可以毫不费力地做到。所以,与其去收集 prompt、琢磨怎么做某件事,不如直接搭一个系统,让 Claude 能从自己的错误中学习,从你的反馈或者从数据中学习,它能去假设、去琢磨出更好的工作方式。对,不做这件事——不组织你的知识、不从错误中学习、把所有东西都装在脑子里——对,这就是 PM 能犯的最大错误。只是来这儿、指望自己能学到更好的 prompt。要分析的数据太多了。Cowork 对比 Claude Code,如果一个 PM 只有一点点时间学其中一个,他应该选哪个,为什么?这不是二选一。我会从 Cowork 开始,因为你在 Cowork 里学到的一切都会帮你更好地理解 code。我从 Cowork 和从 Claude Code 用的是同一个 repo。比如,我刚才已经演示过了,再来一次,我可以选这个 editor 项目。要不把这个去掉。嗯,它有 CLAUDE.md,它会看到同样的文件,它会……所以这就是我在 Visual Studio 里展示过的那个 CLAUDE.md。它会看到同样的文件、同样的结构。所以我可以让它在这个界面里分析推文,我可以在手机上用 dispatch 做,我可以在网页会话、网页 Claude 会话里做,我可以在 Visual Studio 里做——这全都指向同一个 repo。呃,我会从 Cowork 开始,因为这个界面更简单,你不用去适应 explorer 和终端。
[1:22:40] Pavel
This is more user-friendly. You can see files. You can open those files directly here uh without plugins, without how to display markdown, or how how to display HTML. If you see HTML here, you click HTML and it will show you HTML. And and in in Visual Studio, there are certain certain tricks that you you must learn. So, uh I will start with Cogram and just understanding how to work with agents, how to aggregate this knowledge, uh how to define your workflows. And then, once you feel comfortable with Cogram, then add this terminal aspect. That In my opinion, that this is more effective. Should be more effective for many people than trying to uh learn cloud without experiencing Cogram before. What's your hot take on where AI PMs are headed? What does a PM's daily workflow look like in 12 months? I doubt uh that in 12 months uh the role will disappear or something. Uh But, yeah, we are heading into like super individual contributor PM and then CPO CEO at the top. Uh So, I imagine that most of the time you you will be working with agents. Um Orchestrating multiple like multiple agents at the same time. Uh switching switching the context. It will not be easier. It it The work might be even more demanding. Uh But, at the same time, you will focus on There will be less trivial things because those can be automated like writing tickets or debugging something or preparing a presentation. Those are the things that should be automated. So, and eventually people who develop skills from multiple areas. So, like P-shaped or even broader shaped that understand marketing, understand strategy, understand technology, understand products, can talk talk to customers and understand enough to delegate the work assess the the results. What's overhyped versus underhyped in the cloud ecosystem
它更友好。你能看到文件,你可以直接在这里打开那些文件,呃不需要插件,不需要琢磨怎么显示 markdown、怎么显示 HTML。如果你这里看到 HTML,你点一下 HTML,它就会把 HTML 展示给你。而在 Visual Studio 里,有些……有些小窍门是你必须得学的。所以,呃我会从 Cowork 开始,先搞懂怎么跟 agent 协作、怎么把这些知识聚合起来、呃怎么定义你的工作流。然后,等你对 Cowork 上手了,再加上终端这一块。在我看来这样更有效。对很多人来说,应该比那种没体验过 Cowork 就直接去学 Claude(Code)更有效。你对 AI PM 的走向有什么大胆的看法?12 个月后 PM 的日常工作流会是什么样?我不觉得呃 12 个月后这个角色会消失之类的。呃但是,对,我们正走向那种超级个人贡献者型的 PM,然后顶层是 CPO、CEO。呃所以我设想,大部分时间你会跟 agent 一起工作。嗯,同时编排多个……多个 agent。呃,不断切换上下文。这不会变轻松,它……工作可能甚至会更费劲。呃但与此同时,你会专注于……琐碎的事会变少,因为那些可以被自动化,比如写工单、调试什么东西、准备演示文稿——这些都是应该被自动化的。所以,最终能胜出的是那些发展出多个领域技能的人。比如 π 型,甚至更宽的形状——既懂市场、懂战略、懂技术、懂产品,又能跟客户对话,懂得足够多以便去委派工作、评估结果。在 Claude 生态里,什么被过度炒作了,什么又被低估了
[1:24:56] Pavel
right now? I got the impression that people still have not realized
……现在?我有种感觉,大家还没意识到
[1:25:00]
[laughter]
[笑声]
[1:25:01] Pavel
what I just can do especially with the right harness and with the right systems around them. I just started ex- experimenting with that. Like, I've been working with cloud and with agents for 2 years or more, but I just started to like investing heavily in automation and harness myself and I know that I can be much more effective. I doubt there is there is something that is like there's no hype there's not enough hype. So, everything is underhyped right now, guys, in the cloud ecosystem. Go learn what we've showed you right now. Final question for you, Pavel. Your N8N episode that did really well. Is N8N over? I mean, should everything just be done in cloud code now? No.
……我现在到底能做到什么,尤其是配上合适的 harness、配上合适的系统围绕在它周围。我自己也才刚开始实验这些。比如,我跟 Claude、跟 agent 打交道已经 2 年多了,但我才刚开始真正大力投入到自动化和 harness 上,我知道我能变得高效得多。我不觉得有什么是……有什么是被过度炒作的——根本没有炒作,炒作得还不够。所以现在一切都是被低估的,各位,在 Claude 生态里。去学我们刚才给你们展示的东西吧。最后一个问题,Pavel。你那期表现非常好的 n8n 节目——n8n 过时了吗?我是说,现在是不是一切都该在 Claude Code 里做了?不。
[1:25:49]
[laughter]
[笑声]
[1:25:50] Pavel
N8N is still relevant. There are two types of There are two types of automation. One automation is when you automate things for yourself like I want to analyze 100 tweets or I want to draft a a response to customer email then you have this personal automation and you can use cloud code for it. Also, we can automate specific processes inside your code base like code review or release notes or front-end design and have those sub agents, that's fine. Uh but when you want to automate production processes, uh you like the logic that I presented, this is part of the prompt and the agent can respect it, it may not respect it. Uh all we do is edit text files within harness that was defined by Anthropic. We cannot we we cannot tell the agent that, "Hey, if Anthropic API API fails, try three times." Or uh for this, you should always use this tool before uh you should always verify that customer email exists before doing something. Or that customer has access to this data. Uh we just create text files and we hope that agents will follow our instructions. More or less, like uh you cannot hooks, you can add some There are some exceptions from this rule, but overall, we rely on Anthropic harness interpreting our our text files. Uh and then uh this doesn't scale, this is not uh secure enough and this is not effective enough uh for production processes. So, if I want to design uh a system that replies to customer tickets or maybe chats with the customer, I want to have have a hard guidelines, not prompts, that the customer cannot access data of another customer, for example. Or that cannot delete data or um uh if you want to send an email, then yeah, like copy a file that has 1 GB from one place to another. We don't want the agent to do that or maybe even we
n8n 仍然很有用。有两种……有两种类型的自动化。一种是你为自己自动化一些事,比如「我想分析 100 条推文」或者「我想起草一封回复客户邮件的草稿」,这就是个人自动化,你可以用 Claude Code 来做。另外,我们也可以在你的代码库内部自动化一些特定流程,比如 code review、release notes 或者前端设计,让这些 subagent 去做,这没问题。呃但当你想自动化生产流程时——呃比如我刚才展示的那套逻辑,它是 prompt 的一部分,agent 可能遵守它,也可能不遵守。呃我们所做的一切只是在 Anthropic 定义的 harness 里编辑文本文件。我们没法告诉 agent「嘿,如果 Anthropic API 调用失败了,重试三次」。或者呃「做这件事之前你应该总是先用这个工具」「你应该总是先验证客户邮箱存在再做某件事」,或者「客户必须有权访问这些数据」。呃我们只是创建文本文件,然后期望 agent 会遵循我们的指令。多多少少吧,呃你没法……你可以用 hooks 加一些……这条规则有一些例外,但总体上我们依赖 Anthropic 的 harness 去解读我们的文本文件。呃然后呃这不可扩展,这不够安全,对生产流程来说也不够有效。所以,如果我想设计一个呃回复客户工单、或者跟客户聊天的系统,我希望有硬性的准则,而不是 prompt——比如客户不能访问另一个客户的数据,或者不能删除数据,或者嗯呃如果你想发一封邮件……或者比如把一个 1 GB 的文件从一个地方复制到另一个地方,我们不希望 agent 去做这个,甚至我们可能
[1:28:16] Pavel
don't want agent to look at the this data. It should be handled by the code. So, I I I I'm not sure I you remember that we built an agent in three versions and why one was fully autonomous and we started with the least autonomous one where most of the process was code and there was one LLM call. Then we built a hybrid scenario and then the last version was fully autonomous. And in production, everything that doesn't have to be autonomous shouldn't be autonomous. So, we should have code, we should have conditions, we should have guardrails, like if there is a process that should be followed, it should be code. And maybe in this process there are some LLMs LLM calls or agents inside. So, this is this is more difficult to implement, but it is much more cost-effective, it's much more safe than relying on agents respecting instructions. So, when it comes to the takeaway from our n8n episode, as we showed you guys, you want to actually be like less vague. And with these Anthropic-based cloud code systems right now, there's a lot of room for interpretation, less controls. If you're going to build a true production-grade automation for your company, you're still going to be using n8n. You're going to be defining as much as you can some of those hard rules Pavel talked about. Did I summarize it correctly? Yes, and you can also use Anthropic API to to define agents in code. Uh but that's a separate story. And yeah, you can use Anthropic API to define your workflows in code. But this is not the same as organizing text files so that the agent follows the instructions. Um we can also code uh code our workflows with with Anthropic API or uh yeah. OpenAI agentic API or other APIs. Uh the simplest way for a person that doesn't code is using an AI. Makes sense.
……不希望 agent 去看这些数据。这应该由代码来处理。所以我……我不确定你是否还记得,我们用三个版本搭过一个 agent,其中一个是完全自主的,而我们从最不自主的那个开始——大部分流程是代码,只有一次 LLM 调用。然后我们搭了一个混合方案,最后一个版本才是完全自主的。在生产环境里,凡是不需要自主的,就不应该让它自主。所以我们应该有代码、应该有条件判断、应该有护栏——如果有一个必须遵循的流程,它就该是代码。而在这个流程里也许有一些 LLM……LLM 调用或者 agent 嵌在里面。所以这实现起来更难,但它比依赖 agent 遵守指令要划算得多,也安全得多。所以,说到我们那期 n8n 节目的要点,正如我们给你们展示的,你要做到更不模糊一点。而用现在这些基于 Anthropic 的 Claude Code 系统,有很大的解读空间,控制更少。如果你要为你的公司构建一个真正生产级的自动化,你还是会用到 n8n。你会尽可能多地把 Pavel 说的那些硬性规则定义清楚。我总结得对吗?对,而且你也可以用 Anthropic API 在代码里定义 agent。呃但那是另一个话题了。是的,你可以用 Anthropic API 在代码里定义你的工作流。但这跟「组织文本文件让 agent 遵循指令」不是一回事。嗯,我们也可以用 Anthropic API 来写我们的工作流,呃,对。或者用 OpenAI 的 agentic API,或者其他 API。呃,对不写代码的人来说,最简单的方式就是用 AI。有道理。
[1:30:19] Aakash
All right, guys. We have walked you through the AIPM tool universe in today's episode. If you haven't yet, be sure to subscribe to Pavel's newsletter. He has an upcoming Cloudathon starting May 9th, which you may want to participate in. Check out our other episodes if you want to learn more about AI product management and AI or customer discovery. On the May 9th. So, the next in a month we are starting Buildathon, Cloudathon with Cloud and the previous edition it was 250 students that we were building uh real products with an AI and with Lovable and this time we will focus on Cloud Code and an AI at Trigger Dev to build real agentic workflows. Uh so, yeah, I encourage you to check the program. There are only 60 places in total. And you can also see the gallery of the real things that our builders have shipped. So, this is not the theory. There are a few dozens of solutions that are available to browse in the gallery. And you can vote for your favorites and also understand how they were built, what is the architecture, and yeah. Um even download some documents. So, all this is public. And the next class starts next month. All righty, Pavel. Thank you so much for being on the pod. Thank you, Akash. It was a pleasure. I hope you enjoyed that episode. If you could take a moment to double-check that you have followed on Apple and Spotify podcasts, subscribed on YouTube, left a rating or review on Apple or Spotify, and commented on YouTube, all these things will help the algorithm distribute the show to more and more people. As we distribute the show to more people, we can grow the show, improve the quality of the content and the production to get you better insights to stay ahead in your career. Finally, do check out my bundle at bundle.akashsharma.com
好了,各位。今天这一期我们带你们走了一遍 AI PM 的工具宇宙。如果你还没订阅,一定要订阅 Pavel 的 newsletter。他有一个即将开始的 Cloudathon,5 月 9 日开课,你可能会想参加。如果你想了解更多关于 AI 产品管理、AI 或客户洞察的内容,去看看我们的其他节目。在 5 月 9 日。所以,下个月我们要开始 Buildathon、Cloudathon,用 Claude 来做,上一届有 250 名学员,我们用 AI 和 Lovable 一起构建真实的产品,而这一次我们会聚焦在 Claude Code 和 AI,用 Trigger Dev 来构建真正的 agent 工作流。呃,所以,对,我鼓励你去看看这个课程安排。总共只有 60 个名额。你还可以看到我们的学员实际做出来的作品画廊。所以这不是纸上谈兵,画廊里有几十个可以浏览的解决方案,你可以为你喜欢的投票,也可以了解它们是怎么搭出来的、架构是什么样,对。嗯,甚至可以下载一些文档,这些全是公开的。下一期班下个月开课。好啦,Pavel。非常感谢你来上播客。谢谢你,Aakash,很荣幸。希望你喜欢这一期节目。如果你能花点时间确认一下你已经在 Apple 和 Spotify 播客上关注了,在 YouTube 上订阅了,在 Apple 或 Spotify 上留了评分或评价,也在 YouTube 上留了言——所有这些都能帮算法把这档节目分发给越来越多的人。随着节目触达更多人,我们就能把节目做大、提升内容和制作质量,给你更好的洞见,帮你在职业生涯里保持领先。最后,一定去看看我在 bundle.akashsharma.com 上的礼包
[1:32:14] Aakash
to get access to nine AI products for an entire year for free. This includes Dovetail, Mobbin, Linear, Reforge, Build, Descript, and many other amazing tools that will help you as an AI product manager or builder succeed. I'll see you in the next episode.
……免费拿到九款 AI 产品整整一年的使用权。这里面包括 Dovetail、Mobbin、Linear、Reforge、Build、Descript,还有很多其他超棒的工具,会帮你作为 AI 产品经理或构建者取得成功。我们下期节目见。