The Claude Setup That Let a PM Beat 30 Engineering Teams
频道: Aakash Gupta
视频: https://www.youtube.com/watch?v=uEK9ONplfRk
原文语言: en
统计: 共 150 轮 · Jyothi 76 · Aakash 66
[0:00] Jyothi
Understanding which surface to reach for which use case becomes one of the core PM skills that will help you become 10x more effective.
搞清楚什么场景该用哪个界面(surface),会成为 PM 的一项核心技能,能让你的效率提升 10 倍。
[0:08] Aakash
Join Nucle. She's been an AIPM since before it was cool. She's been an AIPM at Netflix, Meta, and Amazon. You posted this on LinkedIn and it caught my eye. You said that you won your internal hackathon [music] against 30 engineering teams and you used this concept of adversarial agents.
欢迎 Jyothi Nookula。早在 AI PM 这个词流行之前她就已经是 AI PM 了——先后在 Netflix、Meta 和 Amazon 做过 AI PM。你在 LinkedIn 上发的那条动态吸引了我:你说你用「对抗性 agent」这个概念,在公司内部 hackathon 上击败了 30 个工程团队。
[0:23] Jyothi
Anthropic had just released a blog post around hardnesses and longunning agents. So I looked into the blog post and they had this concept of adversarial agents. That was what got me [music] the hackathon. Right.
Anthropic 当时刚发布了一篇关于 harness(执行框架)和长时运行 agent 的博客。我研究了那篇文章,里面提到了对抗性 agent 的概念。就是它让我赢下了 hackathon。
[0:34] Aakash
Where does the product manager line end and developer line begin in 2026?
到 2026 年,产品经理的边界在哪里结束,开发者的边界又从哪里开始?
[0:39] Jyothi
Get comfortable web building. Get comfortable with say cloud code with all the cloud ecosystem that we [music] learned today and get comfortable building and putting your ideas out there.
去熟悉网页搭建,去熟悉 Claude Code 和我们今天讲的整个 Claude 生态,习惯于动手构建,把你的想法做出来、发出去。
[0:49] Aakash
How do you use cloud design? How do you build a knowledgebased [music] MCP server for all of your PM context to make cloud 10x more productive? That is what we're going to answer in today's episode.
怎么用 Claude 做设计?怎么为你所有的 PM 上下文搭一个知识库 MCP server,让 Claude 的生产力提升 10 倍?这就是今天这期节目要回答的问题。
[0:59] Jyothi
Now there's this new role coming up called [music] AI builder. Anthropics adopted it. OpenAI's adopted it. Making and [music] building is easy now. Taste is what is important for us to develop.
现在出现了一个新角色,叫 AI builder。Anthropic 采用了这个岗位,OpenAI 也采用了。如今动手做东西已经变得很容易,真正需要我们去培养的是品味(taste)。
[1:10] Aakash
Can you do the big reveal now and help us get that setup going in cloud code?
现在可以来个大揭秘,带我们在 Claude Code 里把那套系统搭起来吗?
[1:14] Jyothi
So here's the thing.
是这样的。
[1:18] Aakash
Before we go any further, do me a favor and check that you are subscribed on YouTube and following on Apple and Spotify podcasts. And if you want to get access to amazing AI tools, check out my bundle where if you become an anal subscriber to my newsletter, you get a full year free of the paid plans of Mobin, Arise, Relay app, Dovetail, Linear, Magic Patterns, Deep Sky, Reforge, Build, Descript, and Speechify. So, be sure to check that out at bundle.ac.com. And now into today's episode. So, I've been thinking about something. We've had advanced tutorials on claude code with analytics on PMOS setup, but how do you actually take the entire cloud ecosystem and make the most out of it from scratch? I keep getting DMs from people who say, "This episode is too complex," or, "I'm not at this level yet. I'm still stuck on Chad GPT." If you're one of those PMs, this episode is going to build you from 0 to 80. We can't get you from 0 to 100 in a single podcast, but we're going to get you the 80% you need to know in 20% of the time. I have brought back Ji Nucla. You guys loved her last episode and specifically the feedback I got was that her structured communication was amazing for beginners. So, she's going to break down for all the beginners how to make the most out of Claude today. Joti, welcome back to the podcast.
在继续之前,帮我个忙:确认你已经在 YouTube 上订阅,并在 Apple 和 Spotify 播客上关注了本节目。如果你想获得一批超棒的 AI 工具,看看我的订阅套餐——成为我 newsletter 的年费订阅者,就能免费获得一整年的付费版 Mobbin、Arize、Relay.app、Dovetail、Linear、Magic Patterns、Deep Sky、Reforge、Build、Descript 和 Speechify。去 bundle.aakashg.com 看看吧。现在进入今天的正题。我一直在想一件事:我们做过 Claude Code 的进阶教程,讲过数据分析,讲过 PM OS 的搭建,但你到底该怎么从零开始,把整个 Claude 生态用到极致?总有人私信我说「这期太复杂了」「我还没到这个水平,还停留在 ChatGPT 阶段」。如果你就是这样的 PM,这期节目会把你从 0 带到 80。一期播客没法把你从 0 带到 100,但我们会用 20% 的时间讲透你需要掌握的那 80%。我请回了 Jyothi Nookula。大家很喜欢她上一期节目,我收到的反馈里特别提到,她结构化的表达方式对新手极其友好。所以今天她要给所有新手拆解,如何把 Claude 用到极致。Jyothi,欢迎回到节目。
[2:44] Jyothi
Super excited to be back. Thank you for having me.
特别高兴能回来,谢谢你的邀请。
[2:47] Aakash
Ji, you posted this on LinkedIn and it caught my eye. You said that you won your internal hackathon against 30 engineering teams and you used this concept of adversarial agents. Can you break down exactly how you won the hackathon?
Jyothi,你在 LinkedIn 上发的那条动态吸引了我。你说你用「对抗性 agent」这个概念,在公司内部 hackathon 上击败了 30 个工程团队。能具体讲讲你是怎么赢下这场 hackathon 的吗?
[3:01] Jyothi
Yes. So a few days before the hackathon um I was trying to see what I could build and Anthropic had just released a blog post around uh hardnesses and longunning agents. So I looked into the blog post and they had this concept of adversarial agents where you build an agent and then you set up configurations in another agent uh telling it what matters most to your company in a way not like eval but more around um um capabilities that you want your um uh agent to test. And so I set up I I started with that idea and then I said let me take this idea. I went into clot code and I was jamming with it for almost a day uh with different configurations and there we go. We I I had an adversarial agent evaluator running. Um it was exactly how I pictured it to be. I even pointed it at our uh company code and integrated that into an actual production running code. Um and that was what got me and my team the uh hackathon uh prize.
好。hackathon 开始前几天,我在琢磨能做点什么,正好 Anthropic 刚发布了一篇关于 harness 和长时运行 agent 的博客。我研究了那篇文章,里面有个对抗性 agent 的概念:你先构建一个 agent,然后在另一个 agent 里做配置,告诉它对你的公司来说什么最重要——不完全是 eval 那种形式,更多是围绕你希望它去测试的能力。我就从这个想法出发,打开 Claude Code,用不同的配置反复折腾了差不多一整天,然后就成了——我跑起来了一个对抗性 agent 评估器,跟我设想的一模一样。我甚至把它指向了我们公司的代码,接进了一段真实在生产环境运行的代码。就是这个东西,让我和我的团队拿下了 hackathon 的奖。
[4:17] Aakash
So that's the promise for you guys. We are going to help you get to that level. Where do we start joti? How can we break this down in a structured way so that people can get to this level at the end of the episode?
这就是我们对各位的承诺:我们会帮你达到这个水平。Jyothi,我们从哪开始?怎么结构化地拆解,让大家在这期结束时能到达这个水平?
[4:30] Jyothi
Great. We'll tackle it today. So, we'll start with understanding the claude stack first and then getting into some of uh the basics um like how do you use cowwork and then getting into cloud code itself. So, let's get started on the claude stack. So, at the bottom of the stack is your models. So claw has haiku sonnet and oppus. They're all different intelligence profiles and which one you need to use when um they have different cost profiles, different intelligence profiles. So that's a decision framework we'll get to in a second. So this is your layer one. On top of your models are what's built your surfaces which use to access these models. So your surfaces could be like cloud um.ai which is on your browser. It could be a desktop app. It could be your mobile app. Um it could be your Chrome plug-in. These are all the interfaces that you use to interact with Claude. Now these are not the same product with just different UIs. they have completely different capabilities and understanding which surface to reach for which use case um becomes one of the core PM skills that will help you become 10x more effective. So this is our second layer. Now on top of this layer is your knowledge base. Now this is where your institutional knowledge lives. Your projects, your skills, your memory, your custom instructions. Now, this is the layer that I think most PMs underinvest in. It's this layer that makes Claude go from being a generic chatbot to actually knowing your context. On that stack is your layer four, which is your integration fabric. For example, your MCPS. Now MCP connects claude to every external system that your organization uses like your Slack, your Google Drive, your Jira, your Salesforce, your internal databases or your own local files. Skills are what extends what your claude knows how to do and what to do with that data. So this is your layer four. And now on top of that is your agents and orchestration. This is where your clawude codework design channels all sit.
好。今天我们这么讲:先理解 Claude 的技术栈,然后过一些基础,比如怎么用 Cowork,再进入 Claude Code 本身。那就从 Claude 栈说起。栈的最底层是模型。Claude 有 Haiku、Sonnet 和 Opus,它们的智能水平不同、成本也不同,什么时候该用哪个,我们等下会讲一个决策框架。这是第一层。模型之上是「界面」(surface),也就是你访问这些模型的入口:可以是浏览器里的 claude.ai,可以是桌面 app、手机 app,也可以是 Chrome 插件。这些都是你和 Claude 交互的界面。但它们不是同一个产品套了不同的 UI——能力完全不同。搞清楚什么场景该用哪个 surface,会成为 PM 的一项核心技能,能让你效率提升 10 倍。这是第二层。再往上一层是知识库,你的机构知识就住在这里:你的 project、skill、记忆、自定义指令。我觉得这是大多数 PM 投入最不足的一层——正是这一层让 Claude 从一个泛泛的聊天机器人,变成真正懂你上下文的助手。再往上是第四层:集成层。比如 MCP——MCP 把 Claude 连接到你组织用的所有外部系统:Slack、Google Drive、Jira、Salesforce、内部数据库,或者你本地的文件。而 skill 扩展的是 Claude「知道怎么做什么」、以及拿到数据后该怎么处理。这是第四层。最顶层是 agent 与编排层,你的 Claude Code、Cowork 这些都在这一层。
[7:05] Aakash
Got it.
明白了。
[7:06] Jyothi
That's how I think about the claude stack.
这就是我理解的 Claude 栈。
[7:09] Aakash
What do people need to know about layers one and two in order to make the most out of the top layers?
关于第一层和第二层,大家需要知道些什么,才能把上面几层用到极致?
[7:14] Jyothi
Yeah. So, let's get into um the models. Now, Haiku is your speed machine. It's the fastest, costefficient, and it's really great for tasks where you need volume over depth. So, let's say you are trying to generate a large number of variance of something or triaging like a pile of documents or you're doing some quick classification or maybe even some tagging. Haiku handles this really well. Now the output won't have the reasoning depth of like your sonnet or your opus but for tasks where depth isn't needed much haiku is sufficient for your use case there. Sonnet is where 90% of my work lives. It has the best quality to cost ratio. So when I'm drafting PRDS or I'm synthesizing user research or I'm doing competitive analysis um or I'm doing stakeholder briefs or I'm thinking about road map I use sonnet sonnet handles all of this extremely well. So opus is for your high stakes high complexity reasoning tasks. So let's say if you're doing some complex trade-off analysis or you're synthesizing genuinely contradictory research or you're doing some long horizon um planning where you need the model like work through second and third order implication. Opus is really good. It has really strong um reasoning capabilities. But I've also noticed from my day-to-day uh working with opus that it also tends to get into this um hallucinated stuck mode a little bit quickly than sonnet where I would use opus and it would get into like one reasoning decision point and it and it would keep um revolving in that local uh maxima and I would have to like literally turn off the chat and move to a new chat and then start all over again to get it out of that thinking mode for example and that's when I sometimes move back to Sonnet because even though it may not have as high a reasoning it's generally a very efficient model to work with and it's also more costefficient than Opus.
好,先说模型。Haiku 是你的速度机器:最快、最省钱,特别适合那些「要量不要深度」的任务。比如你要批量生成某个东西的大量变体、给一堆文档做分流、做快速分类,甚至打标签,Haiku 都处理得很好。它的输出当然不会有 Sonnet 或 Opus 那种推理深度,但对不太需要深度的任务来说,Haiku 完全够用。Sonnet 则承载了我 90% 的工作,它的质量成本比最好。写 PRD、综合用户研究、做竞品分析、写干系人简报、想 roadmap,我都用 Sonnet,它处理这些都非常出色。Opus 留给高风险、高复杂度的推理任务:比如做复杂的权衡分析、综合互相矛盾的研究结论,或者做需要模型推演二阶、三阶影响的长线规划。Opus 在这些方面很强,推理能力非常扎实。不过我日常用 Opus 也发现,它比 Sonnet 更容易陷入一种「幻觉式卡死」的状态——它会在某个推理决策点上打转,困在局部最优里出不来,我只能把对话关掉,开个新对话从头再来,才能把它从那种思维模式里拽出来。这种时候我有时会退回 Sonnet:虽然推理没那么深,但它整体是个非常高效的模型,也比 Opus 省钱。
[9:36] Aakash
Got it. So bring the right model to the right task. It sounds like for 90% of the tasks for PMS you'd recommend Sonic. Yeah, I think that's a good place to start with and then if sonnet doesn't work for the depth that you want, you can always like have open up a chat with Opus and start there.
明白了。所以就是让对的模型干对的活。听起来对 PM 来说,90% 的任务你都推荐 Sonnet。——对,我觉得从 Sonnet 起步很好,如果它的深度不够你要的,随时可以开一个 Opus 的对话接着来。
[9:53] Aakash
What do we need to know about the next layer?
下一层我们需要知道些什么?
[9:55] Jyothi
So, next is your surfaces. Now, clot.ai your which is your web or browser. I think this needs no introduction. Um, everyone's pretty familiar with this. This is where you can use to chat with it. Um the downside is that it doesn't have access to your local system. So if you have some files that you want to access clot.ai may not be able to like directly go and change. Of course you can have like an MCP server but still it's it's I don't prefer it for um for anything that that needs local access. That's when I use desktop. So my cloud core work runs here. It's able to access my files. It's able to um access all the other systems that I have, integrate and and run some scheduled runs, which I'll show you um in a second. I built a podcast guest prep agent in Hyper. The job is simple. Before every interview, give me the guests recent appearances, strongest arguments, company context, sharp question angles, and stuff I should avoid asking [music] because everyone else has already asked. For this run, I pointed it at Howie Lou, CEO and founder of Air Table. The useful part is it can actually go do the research. It's browsed, pulled sources, worked across files and integrations, and then turned the whole thing into a brief I can use before I hit record. Here's the output. Recent [music] appearances, public arguments, company context, question angles, and what not to ask. This is the kind of prep doc I actually want, not a generic summary. It shows what they believe, where their thinking has changed, which questions are obvious, and where the thinking tension might be. Then I saved it as an agent. The output is useful, but the saved agent is the real goal. I don't have to rebuild the whole thing. I point the same agent at a new name, and it already knows the format I like, the sections I care about, and the kind of question framings I come back to. Podcast prep is one example. The bigger idea is recurring work becoming reusable agents. Hyper Agent is built by the team behind Air Table, but it's a separate product. They're offering $1,000 in credits to the first 10,000 subscribers who use my link. Claim yours at hyperagent.com/prouct growth. I also use mobile for when I have a a run kicking off and I can just go for a a walk and I can come back and write while I'm still doing my walk I can look into my phone and see if any of the tasks need my attention. So this has been really helpful um that way. I also use Claude for Chrome plugins especially it's very helpful if you want to do computer use. So, for example, when I'm launching an ad and I want to do some competitive research, I'll kick it off for um through my Claude plugin and it'll use browser use and it will open up a browser. It will do the analysis. It will click through things and say, "Here is what you need to know on how your ad should be against competitors," for example. good for getting into like data that AI agents can't otherwise like LinkedIn or other things like that
接下来是 surface。claude.ai,也就是网页/浏览器版,这个不用多介绍了,大家都很熟,就是用来聊天的地方。缺点是它访问不了你的本地系统——如果你有些本地文件想让它直接处理,claude.ai 做不到直接去改。当然你可以接 MCP server,但凡是需要本地访问的事,我都不太喜欢用它,这时我就用桌面版。我的 Claude Cowork 就跑在桌面版上:它能访问我的文件、连上我其他的系统和集成,还能跑定时任务,等下我演示给你看。(广告)我在 Hyper 里搭了一个播客嘉宾准备 agent,任务很简单:每次采访前,给我这位嘉宾最近的公开露面、最有力的观点、公司背景、犀利的提问角度,以及那些别人都问烂了、我该避开的问题。这次我把它指向了 Airtable 的 CEO 兼创始人 Howie Liu。好用的地方在于它真的会去做研究:浏览网页、拉取来源、跨文件和集成工作,最后整理成一份我录制前真正能用的简报。这是输出:近期露面、公开观点、公司背景、提问角度、以及不该问什么。这才是我真正想要的准备文档,不是泛泛的摘要——它呈现嘉宾相信什么、想法在哪里发生过转变、哪些问题太显而易见、思想张力可能藏在哪。然后我把它存成了一个 agent。输出本身有用,但存下来的 agent 才是真正的目标:我不用每次重建整套流程,把同一个 agent 指向一个新名字,它已经知道我喜欢的格式、我在意的板块、我常用的提问框架。播客准备只是一个例子,更大的想法是:重复性工作变成可复用的 agent。Hyper Agent 出自 Airtable 背后的团队,但是个独立产品。他们为通过我的链接注册的前 10,000 名订阅者提供 1,000 美元额度,去 hyperagent.com/productgrowth 领取。(广告结束)我也用手机版:比如一个任务跑起来之后,我可以出门散个步,边走边在手机上看有没有任务需要我处理。这一点特别方便。我还用 Claude 的 Chrome 插件,尤其适合做 computer use。比如我要投一个广告、想做竞品调研时,就通过 Claude 插件发起,它会用浏览器操作:打开浏览器、做分析、逐个点击页面,然后告诉我「相对竞品,你的广告应该怎么做」。它特别适合获取 AI agent 平时够不着的数据,比如 LinkedIn 之类的。
[13:16] Jyothi
and also good for user testing where you can put up your product up and have um give an instruction to Claude saying go find check check check out this item and you can see how it goes and finds things to see um how well your um product can be understood by agents and where does it fault And uh it also gives you a really good user summary as well if you say behave like a real user and try it and so it'll tell you here are all the things that were confusing. Um and so you can do use it for user testing your products too.
它也很适合做用户测试:把你的产品放上去,给 Claude 一条指令,比如「去找到并查看这个商品」,然后看它怎么一步步找,从而了解 agent 对你产品的理解程度、会在哪里卡壳。你还可以让它「像真实用户一样去试用」,它会告诉你哪些地方让人困惑。所以也可以拿它来给自己的产品做用户测试。
[13:55] Aakash
And are you using cloud code in the desktop app or using it in terminal? Where does that fit in?
那你是在桌面 app 里用 Claude Code,还是在终端里用?它在这个体系里处于什么位置?
[14:00] Jyothi
Oh yeah, that's a good one. I use clot code in IDE because I use clot code to build and so I use cursor or or VS code and today I'll show you with VS code because it's really beginner friendly. So um I use claude code extension in VS code.
哦,这个问题好。我在 IDE 里用 Claude Code,因为我用它来写代码构建东西,所以我用 Cursor 或者 VS Code。今天我会用 VS Code 演示,因为它对新手非常友好——我用的是 VS Code 里的 Claude Code 扩展。
[14:18] Aakash
Is there anything else people need to know about layer 2 or should we move on to layer three?
关于第二层还有什么大家需要知道的吗?还是我们直接进入第三层?
[14:21] Jyothi
Let's move on to layer three. And before we move on to layer three, I'll come back to show you the knowledge base on how to create. But first, let me show you how you you as a PM can 100x your productivity by running a few skills and scheduled runs in co-work.
我们进第三层吧。在讲第三层之前——我稍后会回来演示怎么创建知识库——先让我给你看,作为 PM,怎么靠在 Cowork 里跑几个 skill 和定时任务,把生产力提升 100 倍。
[14:42] Aakash
Awesome.
太好了。
[14:43] Jyothi
Because that will bring us all together on like building your own chief of staff and then I'll show you in plot code how you can um you can do something much more fun. So should people be using chat at all or should they always be using co-work?
因为这会把所有内容串起来——搭建你自己的 chief of staff(幕僚长),之后我再在 Claude Code 里给你看一些更好玩的东西。——那大家还应该用 chat 吗?还是应该一直用 Cowork?
[14:57] Jyothi
So chat is conversational to get you like I have a question what is this versus that or um tell me about a little bit about this information. So it's it's more like a place where you go to search uh instead of going to a Google search. No I just find myself going to clot in chat and asking it some questions. I use co-work for automations. Um, and I'll show you a few today that I use like I have a morning brief. Um, I have my uh G uh Jira uh connected. So I I get my standup brief. So every day it kicks off and tells me here are all the Jira tickets that need your attention and here is um how your project is progressing. Here are four blocked here are three things that have changed. So it gives me my uh brief even before I go to the standup and end of day summary. So there are a lot of things you could do in co-work um in terms of automations to just make your um work life much more easier. Um so you're focusing on things that need the most attention.
chat 是对话式的,用来解决「我有个问题,这个和那个有什么区别」「给我讲讲这个信息」这类需求。它更像一个搜索的去处——现在我不去 Google 搜索了,直接去 Claude 的 chat 里问问题。而 Cowork 我用来做自动化。今天我会演示几个我在用的:比如我有一个晨间简报;我还接了 Jira,所以每天它会自动跑一次,给我一份站会简报,告诉我哪些 Jira 工单需要我关注、项目进展如何——「这里有 4 个被阻塞的,这里有 3 处变更」。这样在我去站会之前就已经拿到了简报,还有每日收尾总结。Cowork 里可以做很多这类自动化,让你的工作轻松很多,把精力集中在最需要关注的事情上。
[16:03] Aakash
Awesome. So can you show us how these work?
太棒了。那你能演示一下这些是怎么用的吗?
[16:05] Jyothi
Sure. So co-work is there on your desktop app. So you need to have your desktop app and you also need to be at least a pro member which is like about $20 per month. So with co-work I can schedule my automations. So you can see I have a few that have scheduled like end of day, daily briefing, daily standup briefing, chief of staff and I'll walk you through each one um right now. So every day at 9:00 a.m. this runs for me where I can say uh and I'll show you a few um as well right now. So you can see my instructions. I'm saying you're my chief of staff. Generate my morning brief for today. Here are your data sources. and I connected it to Google Calendar, Gmail, Google Drive and Jira. How did I do that? Let me show that to you in a second. So, go to customize, go to connectors, and click on the plus. Right now, you can see I have connected to Atlassian Robo, uh, Gmail, Calendar, and Drive. But there's plenty other connectors that you can connect to like Canva, Figma, notion, wherever your data lives. You can connect to it. All that you have to do is just hit a plus and that brings it in and it'll you'll have to authenticate it. Um, and beyond that, that's all you need to do. So, I said here are my data sources. I need you to go into Google calendar, Gmail, Google Drive, and Jira. Pull today's calendar events for each meeting. Capture the title, the time, the attendees, the description, and any attached docs for each meeting with external attendees or something that looks important. Search the Google Drive for any attached doc or recent docs with the meeting title or attendee names. Read enough to know the agenda and search goo uh Gmail for recent threads. Pull Jira items needing my attention. Scan Gmail. You can also add Slack to it and have specific channels that you wanted to review and send it to you as a morning brief. And I said, here's my output format. I want a morning brief, top three things that I need to focus on today. Calendar today. Um, here are the things from inbox that need my attention. Here are the things from Jira that need my attention. And here are some rules. And this is important is I said keep it under 400 words because I don't want to be reading a coffee table edition the first thing in the morning. So keep it under 400 words so it's very easy for me to skim through and understand what I need to focus on what needs my attention immediately. Claude can sometimes pump you up. So I said just give me facts know like great news so don't hype me up. Never invent deadlines or action items. So this is like a guardrail I've put in there and I've asked it to filter aggressively so that I don't have to read anything um or everything all the time and if it's a light day just write a threeline brief and stop. So this is
当然。Cowork 就在你的桌面版 app 里,所以你得装 Claude 桌面 app,而且至少要是 Pro 会员,大概每月 20 美元。有了 Cowork,我就能给自动化任务排时间表。你可以看到我已经排好了几个:end of day(每日收尾)、daily briefing(每日晨报)、daily standup briefing(每日站会简报)、chief of staff(幕僚长),我现在挨个给你讲。每天早上 9 点这个任务会自动帮我跑,我现在也给你演示几个。你能看到我的指令:我说"你是我的 chief of staff,帮我生成今天的晨报,这些是你的数据源"——我把它连到了 Google Calendar、Gmail、Google Drive 和 Jira。怎么连的呢?我马上演示。进入 customize,点 connectors,再点加号。现在你能看到我已经连了 Atlassian(Rovo)、Gmail、Calendar 和 Drive,但还有很多别的 connector 可以连,比如 Canva、Figma、Notion——你的数据在哪儿就连哪儿。只需要点一下加号,它就会接进来,你做一次授权认证就行,除此之外什么都不用干。所以我说:这些是我的数据源,你要去 Google Calendar、Gmail、Google Drive 和 Jira,拉取今天所有的日历日程,每个会议都抓取标题、时间、参会人、描述和附带的文档;对有外部参会者或看起来重要的会议,去 Google Drive 里搜跟会议标题或参会人名字相关的附件或最近的文档,读到能搞清议程为止;再搜 Gmail 里最近的邮件线程;拉取 Jira 里需要我关注的事项;扫一遍 Gmail——你还可以把 Slack 也加进来,指定想让它盯的频道——然后把这些做成一份晨报发给我。我还定义了输出格式:我要一份晨报,先列今天最需要聚焦的三件事,然后是今天的日历、收件箱里需要我处理的事、Jira 里需要我关注的事。另外我加了一些规则,这点很重要:我说全文控制在 400 字以内,因为大清早我可不想读一本茶几画册,400 字以内我扫一眼就知道今天该聚焦什么、什么事需要马上处理。Claude 有时候会给你打鸡血,所以我说只给我事实,别来"好消息!"那一套,别捧我。绝不允许凭空编造 deadline 或行动项——这是我加的护栏。我还要求它狠狠过滤,这样我就不用什么都读;如果哪天事情少,就写个三行简报然后打住。这就是——
[19:04] Aakash
use markdown formatting in order to help it with the headings as well.
你还用了 Markdown 格式,帮它把标题层级也理清楚了。
[19:09] Jyothi
Yes, that makes it easy for clot to read.
对,这样 Claude 读起来更容易。
[19:12] Jyothi
Cool. And so you can see there are a few things that have run previously. So one thing to remember is these automation tasks run only when your laptop is turned on. So if you close your laptop, it doesn't run until your laptop turns back on again. So when you choose the timings, just remember that and so have it at a time when you think your laptop will be on. But otherwise when you turn it on the first time, it will ask you uh and will run that um automation at that time. So like for example, let me show you something that ran. So I ran something that that um that's from May 8th. So it captured a few inbox things that um needs my attention and I can run something now and see how that works. There's a run that started now. So it'll go and collect things and you can see the whole process of how it's thinking. And if you notice I'm using haiko for this. I didn't go and use opers just to like save um some tokens. It asks you for permission. It'll go and pull up things. It'll search email threads. And while that's happening, let me show you the next briefing. So that was my chief of staff morning brief. I also have an end of day which wraps up my day which runs at 5:00 p.m. every day. Now my end of day instructions are very similar. The data sources are similar but the steps are different. So I said read the morning's brief and that's what I had planned to do. Pull what actually happened today like which meetings happened which were cancelled and pull tomorrow's calendar as a preview. And so the output format is like tell me what's shipped, what's slipped, and what's new from today and show me tomorrow at a glance. Again, I have some rules. So this is my instruction for end of day.
好。你可以看到之前已经跑过几次了。有一点要记住:这些自动化任务只在你的笔记本开机时才会运行。如果你合上电脑,它就不跑了,得等电脑重新开机才行。所以选执行时间的时候要注意这一点,挑一个你觉得电脑大概率开着的时间。不过就算错过了,你下次开机时它会问你,然后在那个时候把自动化补跑一遍。举个例子,我给你看一个跑过的结果——这是 5 月 8 日跑的一次,它抓出了收件箱里几件需要我关注的事。我现在也可以现场跑一次给你看效果。你看,一个运行刚刚启动,它会去收集信息,你能看到它思考的完整过程。注意我用的是 Haiku 而不是 Opus,就是为了省点 token。它会向你要权限,然后去拉取信息、搜索邮件线程。趁它在跑,我给你看下一个简报。刚才那个是我的 chief of staff 晨报,我还有一个 end of day,每天下午 5 点运行,帮我收尾一天。end of day 的指令很类似,数据源一样,但步骤不同:我让它先读早上的晨报——那是我今天原本计划要做的——再拉取今天实际发生了什么,比如哪些会开了、哪些取消了,然后把明天的日历拉出来做个预览。输出格式就是:告诉我今天什么交付了、什么延期了、什么是新冒出来的,再让我一眼看到明天的安排。同样,我也加了一些规则。这就是我 end of day 的指令。
[21:15] Aakash
And I guess you could even enhance these if you interested, right? You could probably connect up your analytics. You could add in more context from other systems like your CRM. The limit is just your imagination here. Absolutely. You you can connect it to as many data sources as you want. Um be it even sometimes your um uh Facebook ad systems or your CRM or even your YouTube um and you could get an end of day summary um that captures and it you could also say create a nice dashboard which I'll show you um that I did for Jira where I said the the results uh print it up in a nice dashboard that I can view and it does that for you. And so this is basically taking over a lot of what people would have hired a relay or a lindy last year or a gum loop or a evenmake.com and now you can just build it in claude.
而且我猜你有兴趣的话还能继续增强这些,对吧?可以接上你的分析系统,可以从其他系统比如 CRM 里加更多上下文,上限就是你的想象力。——完全正确。你想连多少数据源都可以,甚至你的 Facebook 广告系统、CRM,甚至 YouTube,然后拿到一份把这些都汇总起来的 end of day 总结。你还可以说"把结果做成一个漂亮的 dashboard"——我等下会给你看我给 Jira 做的那个,我说把结果排版成一个好看的 dashboard,它就真给你做出来。——所以这基本上是把去年大家还要花钱买 Relay、Lindy、Gumloop 甚至 Make.com 才能干的事,现在直接在 Claude 里就能搭出来了。
[22:08] Jyothi
Yes. And one thing it's different from all of those other ones is you would have to like paint um box by box. Think about how the interaction works. Connect each of those and if one thing fails your entire loop fails. that was like how you used to do it before in like say N8N or Lindy or Gumloop or other things that you would want but here you see I'm just giving natural language instruction I can even convert that into a skill um so it's pretty uh robust where it's very easy for me I don't have to think about the architecture I don't have to think how it's connected which box flows into which where is a conditional formatting I don't have to think of any of those
对。而且它跟那些工具有一个本质区别:以前你得一个框一个框地拼,想清楚交互怎么走、把每个节点连起来,只要一个环节挂了,整条流程就全挂了——以前在 n8n、Lindy、Gumloop 那些工具里就是这么干的。但在这儿你看,我只是用自然语言下指令,甚至还能把它转成一个 skill。所以它非常稳健,对我来说特别省心:我不用考虑架构,不用想哪个框连哪个框、条件判断放在哪儿,这些统统不用操心。
[22:52] Aakash
so end of day chief of staff. What are the other two scheduled tasks doing for you?
好,end of day、chief of staff 都讲了。另外两个定时任务是帮你干什么的?
[22:57] Jyothi
So, this one is my standup briefing. This is the one that's connected to Atlassian. Um, that is my Jira Jira board. And so, I said use this Atlassian connector to fetch all issues in the active sprint. And uh here's the brief I wanted done since yesterday, which are the issues moved to done in the last 24 hours. in progress issues, blocked or at risk, new since yesterday, and what's the sprint held? And I said, keep the total under 250 words. And I'll show you an example of this. Just it's asking me for some approval. I approve it, and it's actually rendered it to me really nicely for me to view. And because I asked it to create a dashboard, it's running that. I think what people don't realize is how much better these systems got around December of last year. What really happened that enabled all this to work so much better now?
这个是我的站会简报(standup briefing),就是连着 Atlassian 的那个,也就是我的 Jira 看板。我说:用这个 Atlassian connector 抓取当前 sprint 里的所有 issue,然后按我要的格式出简报——过去 24 小时里哪些 issue 移到了 Done、哪些在进行中、哪些被阻塞或有风险、哪些是昨天之后新增的、sprint 整体健康度如何。我还说全文控制在 250 字以内。我给你看一个例子——它在向我要审批,我批准,然后它就把结果渲染得很漂亮,方便我看。因为我要求它生成一个 dashboard,它现在正在跑。——我觉得很多人没意识到的是,这些系统从去年 12 月左右开始好用了太多。到底发生了什么,让这一切现在能运转得这么好?
[23:52] Jyothi
Improvements in the LLM reasoning capabilities where previously if it I mean previously as well it was much better than what it was 2 years ago. So we're constantly improving but compared to last year the uh new word now is hardness. So the memory, the reasoning capabilities, uh the tools that it can access, all of the underlying systems have improved. And so the latest um improvement is this hardness engineering that is adding so much value into how your systems um behave now.
LLM 推理能力的提升。其实之前也一直在进步——两年前和现在完全没法比,我们是在持续改进的。但跟去年相比,现在的新关键词是 harness(脚手架/工程框架)。记忆、推理能力、它能调用的工具,所有这些底层系统都升级了。而最新的进步就是这种 harness engineering,它给你的系统现在的表现带来了巨大的价值提升。
[24:31] Aakash
And now we have your standup brief. How would you rate this? Is this a good stand-up brief or is this just okay? I think this this is a mocked up one. So therefore, it's showing me a few things which is still a lot better than what I would have had to like go and listen in a call. Um but there's definitely ways I could improve this much more. Like for example, um it's telling me like there's no progress in 24 hours. Um there's I could look at which are which are those ones that have not moved at all and see who is the assigne on those and set up an automation for uh claude to go reach out like ping them on Slack and ask them for an update for example. So like you could set up nested more more automations um as well. So it's it's really helpful to keep away your busy work so you're focusing on actually going and uh solving your customer problems.
现在你的站会简报出来了。你给它打几分?这算一份好的站会简报,还是只能说凑合?——我觉得这份是用 mock 数据演示的,所以它展示的内容有限,但即便如此也比我亲自去听一场站会强多了。当然肯定还有很多可以改进的地方。比如它告诉我"过去 24 小时没有进展"——我就可以去看哪些 issue 完全没动过、看它们的 assignee 是谁,然后再搭一个自动化,让 Claude 主动去 Slack 上 ping 那些人要进展更新。所以你可以往里嵌套更多的自动化。它真的能帮你把杂活都挡掉,让你专注去解决真正的客户问题。
[25:29] Aakash
So if these are the four scheduled tasks, are there any other scheduled tasks that you recommend PMs invest the time in building?
如果说这是你的四个定时任务,那你还推荐 PM 花时间去搭哪些别的定时任务?
[25:36] Jyothi
So what I have here is some examples, but there are lots more you could do. So here's an example. So let's say if there is a ticket a Jira ticket or even a customer support uh ticket that's come from your um from your customers it could automatically be you could create a Jira ticket from it. You could point your clot code to get activated so that it can actually go and implement that and cut a PR and so there's a PR waiting for review.
我这里放的只是几个例子,其实还有很多可以做。举个例子:假如来了一个 Jira 工单,或者客户提了一个客服工单,你可以让它自动生成一张 Jira ticket,再让你的 Claude Code 被触发激活,直接去实现这个需求并提一个 PR——于是就有一个等着你 review 的 PR 了。
[26:07] Aakash
Very cool. So if that's scheduled tasks, I think the next thing you had mentioned this section were skills. What do we need to know about skills? What skills should we have? How do we create them?
非常酷。定时任务讲完了,我记得你这部分接下来要讲的是 skills。关于 skills 我们需要知道什么?该准备哪些 skill?怎么创建?
[26:18] Jyothi
Yes. So if you go to customize again and you can see skills. This is where you can add different skills. I'll show you some examples of some skills. So here's my skill on synthesizing customer interviews. So as PMs we have we sit through lot of customer interviews or at least we get lot of customer interviews for research for feedback for focus group testing for beta testing. We do lots of that and I wanted an easy way for me to have understand what's key what's important and then generate insights from it. So this is my skill that does that which is uh synthesizing customer interviews. Um uh it has like when do you use the skill and um what's the checklist? So it has step by step like inventory the inputs extract observations with citations. Now that's important. I'm not asking to just extract observations. I wanted to site so that it hallucinates less. Use the speaker's own words. Um do not interpret yet. um and separate behavioral observations from stated preferences. I also have additional um MD files listed in here linked so that it could leverage those if needed. Now that's the beauty of skill is skill is not just a markdown file. You also can add functions into it. You can have it link to other skills for example. So what used to happen before a skill was that the whole tool would be loaded into the context and now imagine if you have like 40 tools all of those are loaded into the context. It eats into your context memory. So by default your um your LLM or your model would have very uh limited memory uh for even before you even began asking it anything. What skill does is similar to like progressive disclosure where it add it it it just loads up 50 words of just like the name and the description into the context. Now you can imagine the load is so much lower when the model decides during orchestration based on the question you have asked it goes through the list of tools to see is there a tool that I need to use or is there a skill that I need to use. If it decides that there this skill is valuable based on the description, then it will load the next set of instructions into memory. So that's why skills are powerful because it doesn't eat up or clog your context uh window for your models and it progressively disclosures. And the third is you can link it to more files or more uh skills or functions even like you can have a function where it needs to go run and do something. So
好。还是进入 customize,你能看到 skills,在这里可以添加各种 skill。我给你看几个例子。这是我做的"客户访谈综合分析"skill。作为 PM,我们要参加大量客户访谈,或者至少会收到大量访谈材料——用户研究、反馈收集、焦点小组、beta 测试,做得特别多。我想要一个简便的方式,快速搞清什么是关键、什么重要,然后从中生成洞察。这个 skill 干的就是这个。它里面写了"什么时候用这个 skill",还有一份 checklist,一步一步来:先盘点输入材料,然后"带引用地提取观察"——这点很重要,我不是让它随便提取观察,而是必须附上出处引用,这样它幻觉会更少;要用受访者的原话;这个阶段先不要做解读;还要把行为观察和口头表达的偏好分开。我还在里面链接了额外的 MD 文件,需要时它可以调用。这就是 skill 的妙处:skill 不只是一个 markdown 文件,你还可以往里加函数,还能让它链接到其他 skill。以前没有 skill 的时候,整个工具都会被加载进上下文——想象一下你有 40 个工具,全部塞进上下文里,会疯狂吃掉你的上下文内存。这样一来,你还什么都没问呢,模型可用的记忆就已经所剩无几了。而 skill 做的事类似"渐进式披露"(progressive disclosure):默认只把名称和描述这 50 来个词加载进上下文,负担一下子小了很多。当模型在编排(orchestration)时根据你的问题去扫一遍工具列表,判断"我需要用某个工具吗?需要用某个 skill 吗?",如果它根据描述判断这个 skill 有用,才会把下一层指令加载进内存。这就是 skill 强大的原因:它不会吃掉、堵死你模型的上下文窗口,而是渐进式地披露。第三点是你可以把它链接到更多文件、更多 skill 甚至函数——比如让它去跑一个函数执行某件事。
[29:16] Aakash
I think it was around February of this year when they made skills not just a single markdown file but you could have multiple files and if you aren't using multiple files and your main one isn't less than 500 lines you're really missing out I feel.
我记得是今年 2 月左右,他们把 skill 从单个 markdown 文件升级成可以带多个文件。如果你还没用上多文件,而且主文件也没控制在 500 行以内,我觉得你真的亏大了。
[29:30] Jyothi
Yeah. And so for example I have this um evidence rules.mmd which I'll show it to you in a second. Um so that's in step two. So if step two is invoked then it will go and see evidence underscore rules to ident to understand selection criteria or how to handle ambiguity. Then step three is like cluster into candidate patterns. And look I'm here again linking it to another one called jobs to be done framework. Then I said then apply the pattern threshold. um and then surface the contradictions and then draft hypothesis and then validate every claim against the source codes and that's when I said assemble the final output but I want it in this template and this template is output template so I give it my template so if it gets to step eight is when it will load the output template MD
对。举个例子,我有一个 evidence_rules.md,等下给你看,它挂在第二步里。如果第二步被触发,它就会去读 evidence_rules,搞明白筛选标准、以及遇到模糊情况怎么处理。第三步是把观察聚类成候选模式——你看,我在这里又链接了另一个文件,叫 jobs to be done framework(JTBD 框架)。然后我说应用模式阈值,接着把矛盾点摆出来,再起草假设,然后把每一条论断都对照原始出处做校验。到这时我才说组装最终输出——而且我要求按我的模板来,这个模板就是 output template。我把模板给它,所以只有走到第八步时,它才会加载 output_template.md。
[30:26] Aakash
and right now we're paying a lot of attention to what is actually in the SC skill file. How important is that for PMs versus just letting Claude kind of handle what's in the skill file?
我们现在花了很多心思在 skill 文件里到底写什么上。对 PM 来说,亲自打磨 skill 文件有多重要?还是说交给 Claude 自己看着办就行?
[30:38] Jyothi
So, a lot of times we do use Claude to write the skill file to, but it's also shown, research has shown that um AI generated skill file is less effective than human written skill files. So, that doesn't mean you don't use AI there. Um what I the way I interpret this is put in your human domain knowledge in there to make it work for what you need versus just taking it and automating it from claude and putting it in there. So uh I have used uh Claude a lot to help my skill files write my skill files but then I go and and I add my own tweaks like what's the template that you want how do you want it structured and I work with plot to keep making that changes and from there um add uh and tweak further more um to get to the skill file that I want.
我们确实经常用 Claude 来写 skill 文件,但研究也表明,AI 生成的 skill 文件效果不如人写的。这不是说不能用 AI,我的理解是:要把你自己的领域知识灌进去,让它真正贴合你的需求,而不是让 Claude 自动生成一份就直接拿去用。我自己也大量用 Claude 帮我写 skill 文件,但写完之后我会加上自己的调整——比如你想要什么模板、想怎么组织结构——然后我跟 Claude 一起反复改,在这个基础上继续增补微调,直到得到我想要的那个 skill 文件。
[31:36] Aakash
And how often should we be updating our skill files?
那我们应该多久更新一次 skill 文件?
[31:39] Jyothi
As often as things change for you. So the way to think about skill skills is this is um kind of like a a guide book or a playbook for your claude to know how to do a task for you. So let's say for example PRDS. Now, if your company doesn't change the template of how a BRD is, maybe that's fine, but but your domain may change or your understanding of your domain may continue to change and you do want to like come back and review your skill files. Um maybe say once every quarter depending on how often things change. So the parameters for you to decide is how often does things change in your domain. How frequently do you use that task for? And the third primary thing is how is the output currently? Because if you're not satisfied and you're like it was good but now it doesn't seem to be as good maybe go back to your skill file and say do you need to update it? So it's like that drift as well that you that gives you a cue that you need to go and update it. And what are the most important skill files for PMs to create?
你的情况变多快,就更新多勤。可以这样理解 skill:它是给你的 Claude 用的操作指南或者说 playbook,教它怎么帮你完成一项任务。拿 PRD 举例:如果你公司的 PRD 模板不变,那也许没问题,但你的领域会变,或者你对领域的理解会不断加深,这时你就该回头审视你的 skill 文件——比如每季度一次,视变化频率而定。判断的参数有三个:你所在领域的变化有多频繁;你多久用一次这个任务;第三个也是最重要的——当前的输出质量怎么样。如果你不满意,觉得"以前挺好的,现在好像不如从前了",那就该回到 skill 文件问问自己是不是需要更新了。这种"漂移"本身就是提醒你去更新的信号。——那对 PM 来说,最值得创建的 skill 文件是哪些?
[32:50] Jyothi
Backlog triaging. Give it context. And I'll show you in a second how to do that uh from a context point of view, but give it context. So backlog, um writing PRDs, customer interviews, um even your um support tickets. How do you take a support ticket and how do you put it into a Jira? Right? That could be an automation, but it could also like it could be a skill that is scheduled to run every time there is a a trigger. Now in that case your trigger won't be something that runs time based because there's no like one particular time you're going to get the uh uh support ticket but it could be a trigger when uh this uh whenever there is a support ticket added in your uh service now or Zenesk or wherever your support uh forums are.
Backlog 分诊(triaging)——要给它上下文,等下我会演示从上下文角度怎么做,但一定要给它上下文。所以是:backlog、写 PRD、客户访谈,还有你的客服工单。比如怎么把一个客服工单转成 Jira?这可以是一个自动化,但也可以是一个由触发器驱动、每次触发就运行的 skill。这种情况下你的触发器就不是按时间跑的了——因为客服工单不会在某个固定时间到来——而是每当你的 ServiceNow、Zendesk 或者随便什么客服平台里新增一个工单时触发。
[33:46] Aakash
So is it fair to say you're going to have more skills than scheduled automations? Some of your scheduled automations might reference a skill.
所以可以这么说:你的 skill 会比定时自动化多,而且有些定时自动化会引用某个 skill?
[33:53] Jyothi
That's true. Uh and the way to think about it is most of your scheduled automations are time based. So things that um are more personal productivity based that happen at some sequence like I know I meet my manager once every week. So I know the meeting is always on Wednesday. So I run my um automation on Friday evening to and it maps out saying here are the things that you need to talk to your manager from all these other meetings that you have sat through.
没错。可以这样理解:定时自动化大多是基于时间的,服务那些有固定节奏的个人效率场景。比如我知道我每周和我的 manager 开一次会,会议永远在周三,所以我把自动化安排在周五晚上跑,它会把我这周参加过的所有其他会议梳理一遍,告诉我"这些是你下次需要跟 manager 聊的事"。
[34:22] Aakash
Makes sense. the last layer you talked or I think you were going to show us how to do context in this skill.
有道理。最后一层你提到过——我记得你要演示怎么在这套体系里做 context(上下文)。
[34:28] Jyothi
So here's the thing. So until now what you have done is you've connected it to sources. It can go read all of those sources and u go and do the task for you. But it doesn't learn the people around you. It doesn't learn your connection to people. It doesn't learn it doesn't have that knowledge graph or the knowledge base for what you're working on. And so I wanted to build a chief of staff that understands and is grounded in the knowledge base that I have. Um and so I went to clot code and I said let me spin this up. So I'm going to show you what I'm going to do there. So I'm on VS Code. Now for those um who are looking at this ID for the first time explorer is the place where you can open up your folders and for you to find um uh claude just go into extensions and search for claude code for VS code and um you'll find it there'll be an install just like how you see something else that I haven't installed there'll be an install button that you'll have to click on and that's it. it'll install and then it'll ask you for your login and everything when you um install it. So that way it you're logged in and ready to go always. And it'll show up here um as an icon that you can click and and it'll ask you whether it's a new session or existing session. I'll click on new session and you you'll see how it makes it so much better now that I can just talk to it right here. Of course, I can open the terminal too and it'll I can see if it if I need to like run some commands, but right now I can just talk to it right here in natural language. So, here's my chief of staff um template that I have written where let's say I've joined some company. I'm the senior director there. Um I have I want to build this personal agent that helps me navigate strategy, execution, people, politics. So the agent should learn from my meeting trans transcripts and I use granola for my meeting transcripts. So it should learn from my meeting transcripts. It ingest documents like strategy docs, org charts, PRDs, emails and build a knowledge base over time about people, dynamics, topics, company context. And so I said this is my architecture overview of inputs. Here's my context and my agent. And I said here's my documentation pipeline.
是这样的。到目前为止你做的事情是:把它连上各种数据源,它能去读这些数据源、帮你完成任务。但它不了解你身边的人,不了解你和这些人的关系,它没有那张知识图谱,没有关于你手头工作的知识库。所以我想搭一个真正理解我、扎根在我自己知识库里的 chief of staff。于是我打开 Claude Code,说我们来把这个搭起来。我给你演示我要做什么。我现在在 VS Code 里。给第一次看这个 IDE 的朋友说一下:Explorer 是打开文件夹的地方;要找 Claude,进 Extensions 搜 "Claude Code for VS Code",就会看到一个 Install 按钮——就像你看到我这里其他没装的插件那样——点一下就装好了,装好后它会让你登录,之后你就一直处于登录就绪状态。它会以一个图标的形式出现在这里,点开后它会问你要开新会话还是继续已有会话。我点新会话——你会看到现在体验好太多了,我可以直接在这里跟它对话。当然我也可以打开终端,需要跑命令的时候能看到,但现在我直接在这里用自然语言跟它说话就行。这是我写的 chief of staff 模板:假设我加入了某家公司,是那里的 Senior Director,我想搭一个个人 agent,帮我驾驭战略、执行、人和办公室政治。这个 agent 要从我的会议转写里学习——我用 Granola 做会议转写——所以它要从我的会议记录里学;它要摄入各种文档:战略文档、组织架构图、PRD、邮件,并随时间沉淀出一个关于人、人际动态、话题和公司背景的知识库。然后我写了架构总览:输入是什么、我的 context 和 agent 是什么,还有我的文档处理 pipeline。
[37:08] Aakash
And Claude wrote this, right? Yes, Claude wrote this. Yeah,
这是 Claude 写的,对吧?——对,是 Claude 写的。
[37:12] Aakash
cool.
酷。
[37:13] Jyothi
Um, I told it in natural language like I want XYZ. Here are all the things and it kind of created this whole um MD file that I could use now with Claude again um to build it. So let's say I joined a company as whichever role and I say I want to build a personal AI agent that helps me navigate my work like my strategy execution people and politics. So the agent should like learn from my meeting transcripts. I use granola. You could use zoom. You could use team. You could use whatever you use for transcripts. You just have to like mention that your ingest documents and build a knowledge base over time about people, dynamics, topics, company context. And here's the architecture overview. Now, I gave my use case to Claude and it wrote this up for me and put this architecture overview that I could use it then give it back to Claude again to code it up. And so for part one, there's uh here's my document injection pipeline. So I have like strategy docs what to extract like I want to extract goals priorities metrics timelines and as PMS we are so cross function it's not just our docs we read 50 docs in a week so this is like really helpful for me to like just feed that in and it'll read it up it will store it into a knowledge base and I'll show you that in a second um and it's really cool where the other day I was Um um I was in a meeting. This person was showing me a few things and after the meeting got over and uh we record transcripts um uh in Google Meet and so when the transcript came through my uh chief of staff reviewed it and then it said you know what you should make this person your ally because this person is good at X which you're trying to like get into. And so I'm like oh okay that's great. And then there was something else that um I needed to convey to um somebody and my chief of staff said, "Hey, this is extremely sensitive. Have you thought about XYZ people that you have to inform first before you convey to this person?" And that's so thoughtful because now it it's it understands my org. It understands who is doing what. It understands their personality. So, it's like really powerful. It's like really I have this chief of staff that's telling me always what I need to do. So, here are all the supported documents I wanted to ingest. And here's the document extraction prompt. So, I'm saying you're helping me build a knowledge base about my workplace. I'll share a document. Extract relevant information. Um, and so I said for strategy or planning docs, extract this way. For OGs and OG charts, um, extract these. For PRDS, extract these more. for emails or communication extract these um capabilities. So I have this for each of the ones that I need and I said format the output in this way for it to store in my knowledge base. Um and here's a knowledge base structure. So it has its context K. It has people, topics, meetings, documents, company um and my context like what are my priorities, my OKRs, my preferences, notes, questions, um insights like political landscape and patterns um that it uh identifies or extracts. it can save it here. And these are like patterns observed over time and you can add more as well like to-do for example. It could be a running to-do um that your staff could be maintaining for you. And here are the templates um for the different types. So I said extract this um for it to like save it into the knowledge base. any document that I give uh extract the metadata the summary the key points keep involved what's the relevance to me and my vertical what are the action items and some raw notes um and for people profile again extract these metadata how they operate the communication style meeting behavior what works or doesn't work what they care about what are their motivations so um and what's the relationship to me and Then over time keep reviewing the relationship quality, whether they're a strong ally, friendly, neutral, cautious, friction.
我就是用自然语言告诉它"我要 XYZ,这些是所有要素",它就生成了这份完整的 MD 文件,我可以拿着它再回去找 Claude 把系统搭出来。比如说我以某个职位加入一家公司,我说我要搭一个个人 AI agent,帮我驾驭工作中的战略、执行、人和政治。这个 agent 要从我的会议转写里学习——我用 Granola,你可以用 Zoom、用 Teams,用什么转写工具都行,只要写明就好——它要摄入文档,随时间沉淀出关于人、人际动态、话题、公司背景的知识库。然后是架构总览。我把我的使用场景告诉 Claude,它就写出了这份文档,附上了架构总览,我再把它交回给 Claude 去写代码实现。第一部分是我的文档摄入 pipeline:比如战略文档要提取什么——我要提取目标、优先级、指标、时间线。我们 PM 的工作特别跨职能,读的不只是自己的文档,一周要读 50 份,所以这对我特别有用:喂进去,它读完就存进知识库——等下给你看。而且它真的很神:前几天我在开会,对方给我演示了一些东西,会开完之后——我们在 Google Meet 里录转写——转写一出来,我的 chief of staff 看完就跟我说:"你应该把这个人发展成盟友,因为他擅长 X,而 X 正是你想切入的方向。"我心想,哦,这真不错。还有一次,我有件事需要传达给某人,我的 chief of staff 说:"这件事极其敏感,你有没有想过在告诉这个人之前,得先知会某某某几个人?"这太贴心了——因为它现在理解我的组织,知道谁在做什么,了解每个人的性格。真的很强,就像我身边随时有一位幕僚长在提醒我该做什么。接着是我想摄入的所有支持文档类型,以及文档提取的 prompt。我说:你在帮我建一个关于我工作环境的知识库,我会给你文档,你来提取相关信息。战略或规划类文档按这个方式提取;组织架构图提取这些;PRD 再多提取这些;邮件或沟通类内容提取这些维度。每种类型我都写好了规则,还规定了输出格式,方便它存进知识库。然后是知识库的结构:有 context 目录,下面是 people(人)、topics(话题)、meetings(会议)、documents(文档)、company(公司),还有我自己的 context——我的优先级、OKR、偏好、笔记、疑问,以及它识别提取出的 insights,比如政治格局和长期观察到的模式,都存在这里。你还可以往里加别的,比如 to-do——让你的幕僚长帮你维护一份滚动的待办清单。然后是各类型的模板:任何给它的文档,都要提取元数据、摘要、要点、涉及哪些人、跟我和我负责的方向有什么关联、行动项是什么,再加一些原始笔记。人物画像也一样:提取元数据、他们的做事方式、沟通风格、开会时的表现、什么招对他们管用什么不管用、他们在乎什么、动机是什么、跟我是什么关系。然后随着时间推移持续评估关系质量:是坚定盟友、友好、中立、谨慎,还是有摩擦。
[41:53] Aakash
By the way guys, if you want the exact information that Ji is sharing, you can get all of those in the GitHub link in description,
对了各位,如果你想要 Jyothi 分享的这些完整资料,视频描述里的 GitHub 链接里全都有。
[42:03] Jyothi
organizational dynamics, the observation log, company strategy templates. So when you have your companies sharing you the strategy saying here's what we're going to do in 2026 here are the key things I wanted to extract or structure template and meeting transcript extraction. So when I give it a meeting transcript what I needed to extract um the agent system prompt and this is my uh prompt for the agent uh on you have access to my context KB your job is to help me ramp up fast. Give me strategic advice grounded in context. Help me prepare for meetings. Coach me on people or politics. Help me think through decisions. connect dots across documents and meetings and keep me focused on my priorities. Again, style is like my style, what I like. Don't sugarcoat politics. Um, when I share a document, extra key information, update relevant KD sections. And so, I've given it all of this information, right? So, this is all about like what it needs to do. So, I have this. Now I'm just going to point my claude code to it and say now can you build this knowledge KB and can you put this behind an MCP server so that I can use my cloud desktop to access my knowledge base. So I have an implementation you can see my implementation I'm saying cloud desktop plus MCP. So build MCP servers for KB read and write chat with cloud desktop can also connect to Google Drive, Slack directly. And so I said create the context KB folder structure. Write my goals initial priorities. Ingest any onboarding docs and after your next meeting run the transcripts for extraction and the KB compounds over time. So I'm just going to go to plot code and I'm going to say can you implement? So you're using the at command to pull up that specific file and reference it.
组织动态、观察日志、公司战略模板。比如公司把战略分享给你,说 2026 年我们要做这些事,我要提取哪些关键信息、用什么结构模板;还有会议转写提取——当我给它一份会议 transcript 时,我需要它提取什么。然后是 agent 的 system prompt,这是我给这个 agent 写的提示词:你可以访问我的 context KB,你的任务是帮我快速上手,给我有上下文依据的战略建议,帮我准备会议,在人际和办公室政治上给我教练式指导,帮我想清楚决策,把各个文档和会议之间的线索串起来,并让我聚焦在我的优先事项上。还有风格部分——我的风格、我的偏好,政治问题不要粉饰。当我分享文档时,提取关键信息,更新 KB 里相关的部分。这些信息我都给它了,对吧?这就是它要做的全部事情。有了这个之后,我只要把 Claude Code 指向它,说:现在你能不能把这个知识库搭出来,并把它放到一个 MCP server 后面,让我可以用 Claude Desktop 访问我的知识库。你可以看到我的实现方案写的是 Claude Desktop 加 MCP:为 KB 的读写构建 MCP server,用 Claude Desktop 对话,还可以直接连 Google Drive、Slack。然后我说:创建 context KB 的文件夹结构,写入我的目标和初始优先级,导入所有 onboarding 文档,之后每次开完会就跑一遍 transcript 提取,KB 就会随时间不断累积复利。所以我现在就去 Claude Code 里说:你能实现一下吗?——你是在用 @ 命令调出那个特定文件来引用它,对吧?
[44:07] Jyothi
Yes. So that way it effective. It just knows which one I'm referring to. But even if you don't do it, if you tell it chief of staff agent design, it can go and search through your repository um and find uh the right one for you. So you can implement. Now here I'm if you see what I'm doing um there are there's one thing I want to like show um you is shift tab I can go into plan mode which I can use it for planning again if I do shift tab go into auto mode I can go shift tab ask before edit mode I can go into shift tab again edit automatically where it gets into like actually coding and doing so if you're planning like for example how I planned with it to create that MD file. I was all in that plan mode where I was like, let's just plan. Don't start coding anything. Let's just talk. Um, and now once I'm ready, um, I can shift tab again and go into edit automatically and it will set it off to go do a few things.
对。这样它就明确知道我指的是哪个文件。不过就算你不这么做,只要跟它说 chief of staff agent 的设计文档,它也会自己去搜你的仓库,帮你找到对的那个。所以就让它去实现。这里我想给你们看一个东西:按 shift+tab 我可以进入 plan mode,用来做规划;再按 shift+tab 进入 auto 模式;再按就是 ask before edit 模式;再按一次 shift+tab 就是 edit automatically,也就是真正开始写代码干活的模式。比如做规划的时候——像我之前和它一起规划那个 MD 文件——我全程都在 plan mode 里,就是「我们只做规划,先别写任何代码,就聊」。等我准备好了,再按 shift+tab 切到 edit automatically,它就会自己跑去把事情干完。
[45:18] Aakash
And why do we want this as an MCP server?
那为什么要把它做成一个 MCP server 呢?
[45:22] Jyothi
That's a good call. So um if you want your knowledge base and say Obsidian, you can connect it that way and and put it behind an MCP server and capture it. And here at that point, you just have to say um uh to put this in um Obsidian at that point. But I'm using local. I'm showing it on my local file system because it has a few interesting things. When you're working at a company, you don't want such really personal private data living in some cloud and you want it for example to live on your laptop. So the day when you walk out of the company, the laptop goes to them anyway. So you walk out with no data on your hand. Um, and so I prefer because this is just so much of knowledge base and very private and personal, I keep it on my um, laptop, but you can keep it on Obsidian or Notion or whichever one you want to use for your knowledge base. You just have to change the system prompt at that point.
问得好。如果你想把知识库放在比如 Obsidian 里,可以那样连接,把它放到 MCP server 后面接进来,那时你只要说把内容存进 Obsidian 就行。但我用的是本地——我在本地文件系统上演示,是因为这里有几个有意思的点。在公司工作时,你不希望这么私密的个人数据存在某个云端,你希望它就在你的笔记本上。哪天你离开公司,笔记本反正要交还给公司,你走的时候手上不带任何数据。所以因为这个知识库量很大、又非常私密和个人化,我把它放在自己的笔记本上。但你也可以放在 Obsidian、Notion,或者任何你想用的知识库工具里,只需要相应地改一下 system prompt。
[46:28] Aakash
And what does putting the MCP server on top of the knowledge base help with? Why can't it just be like a set of markdown files and folders?
在知识库上面套一层 MCP server 到底有什么用?为什么不能就是一堆 markdown 文件和文件夹?
[46:36] Jyothi
Yes. So what what it allows it to do is your you can talk to your knowledge base from your desktop app because otherwise where is the knowledge graph sitting? It's sitting in some place and if it's sitting on your computer then it can read and it can write to it. So all those things that we said extract this extract that it will actually go and write it on into your knowledge base automatically. Okay. So it makes it a little bit more portable than a cloud code web session.
好问题。它的作用是让你可以直接从桌面 app 和你的知识库对话。不然的话,这个知识图谱放在哪儿?它就躺在某个地方。而如果它在你的电脑上,Claude 就既能读又能写。我们前面说的那些「提取这个、提取那个」,它会真的自动写进你的知识库里。——明白了,所以它比一个 Claude Code 的网页会话更「可携带」一些。
[47:11] Jyothi
Yes. And so you can go back and even look at all the MD files to see like what it extracted from which um meeting. But over a period of time like right now I have my my knowledge base is like really huge. Um and so I um I don't even go look into the MD files. I just ask desktop um a cloud saying hey um I'm going to meet my manager one-on-one tomorrow. What should I know? And it will go and dig up all the context in the knowledge base and say here are all the things you need to know because it has my todo there. It knows the style of my manager. It was really interesting. It's that this person is a no fuss person and so you should just get to it versus preaming a lot around it
对。你还可以回头翻那些 MD 文件,看它从哪场会议里提取了什么。不过时间久了——像我现在的知识库已经非常庞大了——我根本不去翻 MD 文件了,我就直接问 Claude Desktop:嘿,我明天要和我的经理一对一,我该知道些什么?它就会去知识库里把所有上下文挖出来,告诉我需要知道的所有事情,因为我的 todo 在里面,它也知道我经理的风格。特别有意思的是,它会说这个人是个不喜欢废话的人,所以你应该直奔主题,不要铺垫一大堆。
[47:58] Aakash
because it's capturing across various conversations patterns too.
因为它还在跨越各种对话捕捉行为模式。
[48:03] Aakash
Quality of the data going in is the most important thing. What is the data a PM needs to make sure is hitting their knowledge base?
输入数据的质量是最重要的。一个 PM 需要确保哪些数据进入自己的知识库?
[48:09] Jyothi
Your meeting transcripts for sure because the number of meetings that we attend there's lot of data. there's a lot more richer context there around people their body language when do they push back how do they react um so it's there's lot of like understanding of context that happens there so definitely your meeting transcripts your um key documents that you receive like say strategy docs that I write I I say push this into KB so that it uh remembers so the next time I say I'm working on this project it knows it has context text directly um and any other um documents that you rev you can push that um to your knowledge based tube and then your slack your slack threads that's the other place which is super rich uh beyond meeting transcripts
首先肯定是你的会议 transcript,因为我们开的会太多了,里面数据量很大,而且那里有丰富得多的上下文——关于人、他们的肢体语言、他们什么时候会反对、怎么反应,那里发生着大量对上下文的理解。所以会议 transcript 是必须的。其次是你收到的关键文档,比如战略文档——我自己写的战略文档,我会说「把这个推进 KB」,这样它就记住了,下次我说我在做这个项目时,它直接就有上下文。其他任何你评审过的文档也都可以推进知识库。然后是你的 Slack——Slack 上的讨论串,那是会议 transcript 之外另一个信息极其丰富的地方。
[49:02] Aakash
and so do you need to like update your KB somehow or do you set a scheduled task to update your KB or how do you make sure that it's kind of not
那你需要手动更新 KB 吗?还是设置一个定时任务去更新?你怎么保证它不会过时?
[49:09] Jyothi
no so every time every time you have a meeting transcript it writes to the KB
不需要。每一次有会议 transcript 进来,它都会自动写进 KB。
[49:15] Aakash
and how do you set that up?
这个是怎么设置的?
[49:17] Jyothi
Yeah, I'll just show that once this is done. So, uh it's again your MCP um so if it is let's say um you have granola so every time or you have Google meet and every time there is a new meet recording that hits or a transcript that hits your inbox um you could set up a co-work um automation to say use this um and update KB for example. So set up some sort of automation to make your KB updating. Make your KB an MCP server so that you can access it from regular cloud chats, not just cloud code web sessions.
好,等这个跑完我就演示。还是靠你的 MCP——比如说你用 Granola,或者用 Google Meet,每次有新的会议录音或者 transcript 到达你的收件箱,你就可以设置一个 Cowork 自动化,说「用这个去更新 KB」。总之就是设置某种自动化让你的 KB 保持更新。把你的 KB 做成 MCP server,这样你在普通的 Claude 对话里就能访问它,而不只是在 Claude Code 的网页会话里。
[49:52] Aakash
And then you're really putting everything together within layer 3. You've got skills, you've got memory. Is there anything people need to know around projects?
然后在第 3 层里你其实是把所有东西整合到一起了——有 skills、有 memory。关于 projects,大家还需要了解什么吗?
[50:03]
Yes. Um in a quick second once this is done, it'll actually ask me to create a project and put the instructions in there. H okay. And what are the projects PMS should be creating?
有的。稍等一下,等这个跑完,它其实会让我创建一个 project,把指令放进去。——嗯,好。那 PM 们应该创建哪些 projects 呢?
[50:16] Jyothi
The way to think about it is organize it like your folders which have unique information um related to it. So let's say you have you work on say three projects uh at company like say you're you're PMing three swim lanes and each swim lane could be a project and you can have the necessary context um that you need in there um in as project instructions that you could then use um um for your claude to understand that a little better.
可以这样理解:像整理文件夹一样去组织,每个文件夹里放与它相关的独特信息。比如说你在公司同时负责三个项目——你在 PM 三条业务线——那每条线都可以是一个 project,你把需要的上下文以 project instructions 的形式放进去,这样你的 Claude 就能更好地理解那条线。
[50:50] Aakash
Got it. Let's do it.
明白了。我们来试试吧。
[50:52] Jyothi
So here it's done. There's a knowledge base at context KB full structure. So I'll just show that to you. And it's also MCP server is at the server.py installed here. It exposes these tools. Um and it appended chief of staff server uh alongside my existing file system server. And so to activate, I just need to fully quit my desktop and reopen the cloud desktop and reopen and then chat uh start a new chat, insert the chief of staff um system prompt from the slash menu and then try list everything in my KB or paste a gran granola transcript and run extract meeting. Um and it al it also added um details and troubleshooting in a readme um as well. So let me quickly pull up um my context KB and just show you how that look and let's say I don't know where that is for example I could also ask it where is it for but in this case I will pull it up and show you can see created two folders context KB and my MCP server I hit on context KB it has created these folders in a nice way company documents insights meetings all of the things that I asked it to like capture It has ready folders and so you can see this MD file setup for everything. So as and how it's extracting it will write into these MD files. And this is the MCP server. So let's see what it has asked me to do from the slash menu. Okay. So let me first quit my desktop app. Quit is just command Q. So I quit it and then I reopen. So I have reopened. Now, if I go into customize and connectors, let's see. You can see it has installed my chief of staff local MCP server.
好,这边跑完了。知识库已经在 context KB 下建好了完整结构,我展示给你看。MCP server 也装在这里的 server.py 里,它暴露了这些工具,并且把 chief of staff server 追加到了我已有的文件系统 server 旁边。要激活的话,我只需要完全退出 Claude Desktop 再重新打开,然后开一个新对话,从斜杠菜单里插入 chief of staff 的 system prompt,然后试试「列出我 KB 里的所有内容」,或者粘贴一份 Granola 的 transcript 然后跑 extract meeting。它还在 README 里写了详细说明和故障排查。让我快速打开我的 context KB 给你看看长什么样——假如我不知道它在哪儿,我也可以直接问它在哪,不过这次我直接打开。你可以看到它创建了两个文件夹:context KB 和我的 MCP server。点进 context KB,它把文件夹建得很漂亮:公司文档、洞察、会议……所有我要求它捕捉的东西都有现成的文件夹,每一项都建好了对应的 MD 文件。之后它每提取一次,就会写进这些 MD 文件。这个是 MCP server。我们看看它让我在斜杠菜单里做什么。好,我先退出桌面 app——Quit 就是 Command+Q——退出后重新打开。现在打开了,进入 customize 和 connectors 看一下,可以看到它已经装好了我的 chief of staff 本地 MCP server。
[52:51] Aakash
This is so cool. We've done a lot of cloud guides and nobody has really shown this feature before.
这太酷了。我们做过很多 Claude 教程,还没人真正展示过这个功能。
[52:56] Jyothi
This has saved me so much of time. Like, it's it's literally my productivity booster and it tells me things and nuances that I might have forgotten otherwise. Okay, so we restarted um insert the chief of staff system prompt from /menu. Now I could say or let me say where is it? Where is the um system prompt? Let's say I don't know right could just ask it. So look it's given me the system prom lives inside the MCP server here it's the chief of staff system okay how to use it. So after you restart I can in a new chat type slash and you'll see the system prompts. Okay. So let's go here. So after I restarted I'll create a new chat and I will say what did it ask me to do or it's saying you can use this include project custom instructions feed. So I'll go create a project and put that as instructions. So, let's say I'm going to say this is um I'm going to create a project and I'm just going to use um say I'm working on a product called meal planner and all my meetings or it could also be company X at like Uber level if you just want it to be like one. I can just put company X and here's where I I I can give um my instructions. Let me first like create the project and then I can add my instructions here from so I can go into my file system in my so let's say I'm not able to find it I'll say can you create the system prompt as a MD file that I can paste paste into project instructions. So, it's writing the prompt. So, there you go. Here's my prompt. I can just copy it and we'll I'll show you what has. So, I'm going back to my project instructions. I'm just going to paste this. So, it's saying here, you're my personal chief of staff, an AI advisor who helps me navigate. I'm so and so. I just started. You have MCP access. Use these tools. Your job is to wrap me up fast. Here's my style. When I share a document, when I share a meeting transcript. So, we do. You can modify this more and refine this more, but for now, I'm fine with this. So, I save that instruction. Now, if I give it a meeting transcript, let's see what it does. Let's say I have this interview that I got. I'm going to add this here and I'll say log it into KB. Let's see what it does. I'll log this interview into KB.
这为我省了太多时间了,它简直就是我的生产力加速器,它会告诉我那些我自己可能会忘掉的细节和微妙之处。好,重启完了,从斜杠菜单插入 chief of staff 的 system prompt。假设我不知道 system prompt 在哪儿,可以直接问它。看,它告诉我 system prompt 就在 MCP server 里,是 chief of staff system,还有使用方法:重启之后在新对话里输入斜杠,就能看到 system prompt。好,重启之后我新建一个对话,看它让我做什么——它说你可以用 project 的自定义指令功能。那我就去创建一个 project,把它作为 instructions 放进去。比如说我在做一个叫 meal planner 的产品,所有会议都归到这里;也可以是「公司 X」这种 Uber 级别的,如果你只想建一个的话。我就填 company X,然后在这里写我的指令。我先把 project 建好,再从文件系统里把指令加进来。假设我找不到那个文件,我就说:你能把 system prompt 生成一个 MD 文件,让我可以粘贴进 project instructions 吗?它正在写这个 prompt——好,出来了,这是我的 prompt,直接复制。回到我的 project instructions,粘贴进去。它写的是:你是我的个人 chief of staff,一个帮我导航的 AI 顾问;我是某某某,刚入职;你有 MCP 访问权限,使用这些工具;你的职责是帮我快速上手;这是我的风格;当我分享文档时怎样、分享会议 transcript 时怎样。你可以再修改和打磨,但现在这样我就够用了。保存指令。现在如果我给它一份会议 transcript,看它会做什么。比如我手上有这份访谈记录,我把它加进来,然后说:把它记录进 KB。看看它怎么做——把这份访谈 log 进 KB。
[56:14] Jyothi
So, ideally, it should be using kind of the right tool in our KBMCP server. So it wants to use my KB. So it'll ask for permission once because it's the first time you have set up it's asking all the permissions but after that it's pretty smooth.
理想情况下,它应该会调用我们 KB MCP server 里对的那个工具。你看它想用我的 KB。它会请求一次权限,因为是第一次设置,所有权限都会问一遍,但之后就很顺畅了。
[56:29] Aakash
Is there like an always allow permissions mode on the app? Like there is dangerously skip permissions in cloud code
这个 app 里有没有那种「始终允许」的权限模式?就像 Claude Code 里的 dangerously skip permissions 那样。
[56:36] Jyothi
there. There is but the thing is the first time it will still ask for it um because it's asking it's accessing your tools and I do give it always allow. So then it'll run. The next time I send it something, it doesn't ask for permission. It'll just go directly read. So you can see it's loaded a bunch of um MCP tools. Um and so it's going and saving that in meetings template.md and then it'll give you some something around one discipline node is this is a single interview. So there's no pattern yet. But if I add like a few more um it'll generate some patterns and I and you can actually just have like a co-work either. So there are a couple of options, right? So you can every meeting transcript you can just paste into this and it will automatically extract and fill your knowledge base or you can have a co-work um automation that every time there is a meeting transcript in your email um and look for what how the uh meeting transcript lands like Google meet has like a Google meet meet recordings or some some transcript words. So use that and say every time this lands in my inbox automatically uh log this into my KB and it will do this all this thing automatically for you. And if you're an email heavy company you could say every email that I get just log it into my KB. It will do that to you.
有的。不过第一次它还是会问,因为它在访问你的工具,我一般都选「始终允许」。之后它就会直接跑——下次我再发东西给它,它不会再要权限,直接就去读了。你可以看到它加载了一堆 MCP 工具,正在把内容存进 meetings template.md,然后它会给你一个提示:这只是单个访谈,还没有形成模式。但如果我再加几份,它就会归纳出一些模式。而且你完全可以配一个 Cowork 自动化。所以有两个选项:一是每份会议 transcript 你都手动粘贴进来,它自动提取并填充你的知识库;二是设一个 Cowork 自动化——每当你的邮箱里出现会议 transcript(可以看一下 transcript 是怎么到达的,比如 Google Meet 会有 meet recordings 或者带 transcript 字样的邮件),就用那个特征说「每次这类邮件进入我的收件箱,就自动把它 log 进我的 KB」,它会全自动帮你完成。如果你的公司特别依赖邮件,你甚至可以说「我收到的每封邮件都 log 进 KB」,它也能做到。
[58:05] Aakash
Oh man, you might burn some tokens that way.
老兄,那样你可要烧不少 token 了。
[58:08] Jyothi
You will but you have such a rich knowledge base at that point in time where it will connect all the pieces together.
确实会,但那时你会拥有一个极其丰富的知识库,它能把所有碎片串联起来。
[58:15] Jyothi
So you can see now both files are now in KB. Here's the things that's logged. um a couple of judgment calls. So you see this is the first time so it's not like it's going to give you the best of insight but you can see it's giving me things worth my attention. So I'm a senior director um um the AI coach is part of Lumen in your lane and it's the weakest thing in this interview. Maya whoever is this uh user ignores it and two times it showed up that she bounced off it. I just wanted a yes or a no. Um, and so that's a clean agent UX signal where and it's this kind of thread worth watching as more interviews come in. So you see how it's it's just one interview in, but it's giving you insights and things that you need to watch for.
你可以看到,现在两个文件都进了 KB。这是记录下来的内容,还有几个判断性结论。这是第一次跑,所以它不会给你最深刻的洞察,但你能看到它已经在给我值得注意的东西了。我是资深总监,AI coach 属于你负责的 Lumen 业务线,而它是这次访谈里表现最弱的部分。Maya——也就是这位用户——直接无视了它,有两次都表现出她从它那里弹开了:「我只想要一个是或否的回答」。这是一个很干净的 agent UX 信号,是那种随着更多访谈进来值得持续关注的线索。你看,虽然才一份访谈,它已经在给你洞察和需要盯的点了。
[59:02] Aakash
Love it. Okay, shall we move on to layer 4? I feel like we've already a little bit talked about layer 4 with people because we've showed them an MCP and layer 4's integrations, but you had this really cool LinkedIn post which maybe you can teach us a little bit about right now. Um, what what exactly do people need to know about MCPs?
太棒了。好,我们进入第 4 层吧?其实我们已经稍微聊到第 4 层了,因为我们给大家展示了一个 MCP,而第 4 层就是集成。不过你之前发过一篇特别棒的 LinkedIn 帖子,也许现在可以给我们讲讲。关于 MCP,大家到底需要知道什么?
[59:24] Jyothi
Yes, so MCP is the way that allows you to connect to different capabilities like your Gmail, Slack. I think we connected to a bunch uh in our cowwork. Um and so I didn't I showed you two things. I showed you remote MCP. I also showed you local MCP like your knowledge KBE is your local MCP that uh you're accessing.
好,MCP 是让你连接各种能力的方式,比如你的 Gmail、Slack。我记得我们在 Cowork 里连了一堆。我给你们展示了两种:一种是远程 MCP,另一种是本地 MCP——比如你的知识库 KB 就是你正在访问的本地 MCP。
[59:49] Aakash
What integrations or MCPs do PMS need to make sure that they have?
PM 们需要确保自己配上哪些集成或者 MCP?
[59:56] Jyothi
So look at the tools that you use more often. So like Gmail uh for example assuming your company uses Gmail for emails you want to connect that calendar you want to connect that you want to connect your Slack you want to connect your uh meeting transcripts wherever they are stored like if that's granola or if that's um um Google meet or zoom recordings you want to connect those you want to connect your CRM your dashboards um your um uh Jira boards your um maybe you're using amp Amplitude for analytics, connect that there. Uh maybe you're using some other tool like radar for observability to monitor your the performance of your application. Connect it. Um you it's the the possibilities are really endless. Um so for example one of the tool that I had connected at work um is uh Nvidia biono model uh to um help uh show uh and do a drug prediction based on a few uh component libraries. So it's like really like it's you're only limited by what you can imagine. But that doesn't mean you go you just go on a MCP shopping spree. So, I would say start off with like connecting what works for you and what use cases you're trying to solve. And so, if you're a beginner, try to follow through this video and do some of the initial automations that I showed you in cowwork to just get started, get your hands uh wet and then go build this chief of staff for yourself. And you could ask your chief of staff, what else should I connect to? and it will tell you here are the list of servers that you need to connect to because I'm seeing this being mentioned in meetings and you don't have the access to that.
所以先看看你平时用得最多的工具。比如 Gmail——假设你们公司用 Gmail 收发邮件,那你就把它连上;日历要连上;Slack 要连上;还有你的会议转录,不管存在哪里——是 Granola 也好,Google Meet 或 Zoom 的录音也好,都连上;再连上你的 CRM、你的数据看板、你的 Jira 看板;如果你用 Amplitude 做分析,也把它接进来;也许你还在用 Radar 之类的可观测性工具来监控应用性能,也接上。可能性真的是无穷无尽的。举个例子,我在公司接过的一个工具是 NVIDIA 的 BioNeMo 模型,用来基于几个组件库做药物预测展示。所以说,唯一限制你的就是你的想象力。但这不代表你要去搞一场 MCP 疯狂采购。我的建议是,从对你真正有用的、你想解决的场景开始连。如果你是新手,可以跟着这期视频,先把我在 Cowork 里演示的那几个初级自动化做一遍,先上手练练,然后再去给自己搭这个 chief of staff。之后你可以直接问你的 chief of staff:我还应该连什么?它会告诉你:这里有一串你需要接的 server,因为我看到会议里反复提到这些东西,而你还没有访问权限。
[1:01:44] Aakash
Love it. So, you can actually progressively build on your connections with your chief of staff. Start with that meeting transcript. Let's move into layer five, shall we?
太棒了。所以你其实可以跟着你的 chief of staff 一步步渐进式地扩展连接,先从会议转录开始。那我们进入第五层吧?
[1:01:53] Jyothi
Yeah.
好。
[1:01:54] Aakash
Lovely. Perfect.
很好,完美。
[1:01:55] Jyothi
Okay.
OK。
[1:01:56]
All right. So, that covers layer four. We've now done layer 1, two, three, four. The next is five. What do people need to know about agents and agent harnesses? So we have built clot we have used clot code for building our capabilities. We have used co-work which are all sitting in your five layer five. Now I want to show you design cla. So the thing with claude design is you have to do claude.ai/design. Um so let me show you that claude.ai/design. It is not integrated directly in your claude.ai yet. You have to go through claude. Oh,
好,第四层就讲完了。我们已经过完了第一、二、三、四层,接下来是第五层。关于 agent 和 agent harness,大家需要知道什么?——我们已经用 Claude Code 搭建了各种能力,也用了 Cowork,这些都属于你的第五层。现在我想给你们展示 Claude Design。用 Claude Design 有个点要注意:你得直接访问 claude.ai/design。我来演示一下——claude.ai/design。它现在还没有直接集成进 claude.ai 主界面,你得走这个入口。哦——
[1:02:35] Jyothi
there we go.
好了,出来了。
[1:02:36] Jyothi
And so it's you can see it's in research preview. Now you can as PMS we do a lot of um design work. We prototype, we create slide decks. Um we create mock applications and plot design really works with a lot of those things. So for example, you can I'll show you a few things here. So prototype I can give it a name. You can see there is wireframe and high fidelity. So you can choose which type you want on slide deck. Uh you can give the project nail and you can even attach um your um speaker notes. Um and it'll create a deck. Again, you can use an animationbased um template to create uh something. And you have an other. Now on your right, you'll see you have recent any designs that you have worked. There are examples um that you can use to get inspiration and you can use as templates. And there's something called design systems. Now design system is something interesting. Now, if you if you have a brand um color, like for example, companies, they have a design guide, um you would want your slides or your uh wireframes or your markups to look similar to what your console is or what your company's colors are. And then you can use this design guide here. You can just click on create. You can um link you can either give it a link on GitHub or you can upload a Figma file or you can add all your assets here and create a design system. I'll show you an example of a LinkedIn post I did. I just gave it my post and I said can you create visuals for it? Um and and so it created this kusal that I wanted in the colors of my product nextg product manager. Um so um it created this eight card corrosal based on the uh text I gave it. So I gave it my post my LinkedIn post. I said here is what my LinkedIn post is about. Can you create this? And created this for me. Now here are some cool things I want to show you. Now it's built this. Now let's say I want to mark it up. I want to tell Claude to change something. So maybe say I wanted to tell uh make layers and use orange highlight color and Claude can go and change just this one piece.
你可以看到它还在 research preview 阶段。作为 PM,我们要做大量的设计类工作:做原型、做幻灯片、做 mock 应用,而 Claude Design 跟这些事都很搭。我给你们演示几个功能。比如原型(prototype),我可以给它起个名字,你可以看到有 wireframe(线框图)和 high fidelity(高保真)两种,可以自己选类型。幻灯片(slide deck)这边,你可以填项目名称,甚至可以把演讲者备注附上去,它就会生成一套 deck。你还可以用带动画的模板来做东西,另外还有一个「其他」选项。右边你会看到「最近」,也就是你做过的所有设计;还有一些示例,可以拿来找灵感或当模板用。还有一个东西叫 design system(设计系统),这个很有意思。如果你有品牌色——比如公司都有自己的设计规范——你会希望你的幻灯片、线框图、标注稿看起来跟你们的产品控制台、跟公司的品牌色一致。这时候就可以用这里的设计规范:点创建,可以给它一个 GitHub 链接,或者上传 Figma 文件,或者把你所有的素材传上来,生成一套设计系统。我给你们看一个我做过的 LinkedIn 帖子的例子:我把帖子原文丢给它,说「能帮我做配图吗」,它就按我的产品 Next Gen Product Manager 的品牌配色,做出了我想要的这套轮播图(carousel)——基于我给的文字生成了这套八张卡片的 carousel。我就是把 LinkedIn 帖子给它,说「我的帖子讲的是这个,能做一套吗」,它就做出来了。接下来给你们看几个很酷的功能。它已经生成好了,假设我现在想在上面做标注,让 Claude 改点东西——比如我想说「把 layers 这块改用橙色高亮」,Claude 就能只改这一处。
[1:05:24] Aakash
So it's got that visual editor built in now.
所以它现在内置了可视化编辑器。
[1:05:27] Jyothi
Yeah. And you can also drop things. That's pretty interesting. Um for uh uh for editing. So I can edit I can I can give it instructions right here and say edit this. I can leave comments.
对。而且你还可以直接拖东西进去,这挺有意思的。编辑的话,我可以直接在这里给它下指令说「改这个」,也可以留评论。
[1:05:44] Jyothi
Yeah I can change I can I can do comments like
对,我可以改,也可以像这样加评论——
[1:05:48] Aakash
oh like we used to do in Figma but now the will execute the edit.
哦,就像我们以前在 Figma 里做的那样,只不过现在它会真的把修改执行掉。
[1:05:52] Jyothi
Yeah. So I can give it comments right here and send it to Claude. Um, and I can um even drag uh and I can like just draw and say make it counter clockwise.
对。我可以直接在这里写评论发给 Claude。我甚至可以拖拽,或者直接画一笔,然后说「把它改成逆时针」。
[1:06:08] Aakash
Wow. So, this is a carousel, but should PMS basically be creating all their presentations in Claude Design now?
哇。这是个轮播图的例子,那 PM 们现在是不是干脆应该把所有演示文稿都放到 Claude Design 里做?
[1:06:14] Jyothi
I used to be a big GMA fan and now I just use Cloud Design for everything. It Wow, it consumes more tokens. The token budget is different. Um, but it's been very rewarding where I don't have to sit and create slide decks anymore and does it in my brand guide. Um, and so it just doesn't even look any different. So um it was uh funny. I had a uh meeting with my CEO um and 1 hour before I I have my content. I pushed it to Claude. I made it create a slide deck. It's like looks so um professional. It doesn't look like it was just done like a few uh minutes before. It looks like I spent several hours to sit and create it.
我以前是 Gamma 的忠实用户,现在什么都用 Claude Design 做。它确实更耗 token,token 预算不太一样,但回报非常大——我再也不用坐在那儿手搓幻灯片了,而且它是按我的品牌规范做的,看起来跟以前完全没差别。说个好玩的事:有次我要跟 CEO 开会,开会前一小时我才把内容准备好,丢给 Claude 让它生成一套幻灯片,结果看起来特别专业,完全不像是几分钟前才赶出来的,倒像是我花了好几个小时精心做的。
[1:07:00] Aakash
Wow. So, it created a CEO level presentation for you in an hour that looked like it took hours. Very cool. Should PMS be making prototypes in cloud design?
哇。所以它一小时内就给你做出了一套 CEO 级别的演示文稿,看起来却像花了好几个小时。太酷了。那 PM 应该用 Claude Design 做原型吗?
[1:07:12] Jyothi
So, here's the thing. So, you there are different types when you need different uh levels of uh prototypes. So for something quick and dirty where you're like is this how I want this to be um use um use clot design if you need but I am more a clot code user because I'm like I'll just spin up and go create that app really quickly and I I it's something like an app so people could like click on it and see how it works and you get really good feedback that way. But you also could use this to create your slide decks to make presentations to your company about what's the feedback from that user interview. You could use this to create quick uh um design patterns um that you want to like share because your app may be like easy for user testing. But maybe you you want this to like um uh create some marketing content to share with your uh marketing team or uh you want to create training uh uh playbooks for your sales and accounts team um to tell how to how to go use this product. And these are like really quick designs you can generate uh without have which looks polished and professional for them to like just go put it into the um into wherever your uh uh knowledge base is.
是这样的:不同场景需要不同层次的原型。如果是那种快糙猛的验证——「我想要的是不是这个样子」——需要的话可以用 Claude Design。但我个人更偏向用 Claude Code,因为我会直接快速起一个应用,做成大家真的能点、能操作的东西,这样拿到的反馈质量会高很多。不过你也可以用 Claude Design 做幻灯片,向公司汇报用户访谈的反馈;可以用它快速产出你想分享的设计样式;也许你想做一些营销素材给市场团队,或者给销售和客户团队做产品使用的培训 playbook。这些都是能很快生成、又显得精致专业的设计,他们直接放进你们的知识库就能用。
[1:08:37] Aakash
Awesome. You mentioned you use cloud code for prototyping and that's where I wanted to take it next. So we started this episode with adversarial agents and your cloud code setup to win the hackathon. Can you do the big reveal now and help us get that setup going in Cloud Code?
太棒了。你提到你用 Claude Code 做原型,我接下来正想聊这个。这期节目开头我们就讲了对抗式 agent,还有你靠 Claude Code 的那套配置赢下了 hackathon。现在能不能来个大揭秘,带我们在 Claude Code 里把那套东西搭起来?
[1:08:53] Jyothi
So, here's the thing. I have to create that whole thing here.
这样的话,我得在这里把整套东西从头搭一遍。
[1:08:56] Aakash
All right, let's do it.
好啊,那就搭吧。
[1:08:58] Aakash
What do you have? We have the time.
有什么问题?我们时间够。
[1:09:00] Jyothi
Let's do it.
那就开始。
[1:09:01] Jyothi
So, I'll create a new session. Let me actually open up a new project. So, you can see I'm opening up a new window so my previous one doesn't interfere with this. And I'll go create a folder for this. So, I can open it up. So now let me open and you see it's a clean folder. I'm just going to start a new session here and I'm going to say let's build an adversarial evaluator.
我先新建一个会话。我干脆开一个新项目——你可以看到我开了一个新窗口,这样之前的东西不会干扰这次演示。我先建一个文件夹,然后打开它。你看,这是个干干净净的空文件夹。我在这里启动一个新会话,然后输入:我们来构建一个对抗式评估器(adversarial evaluator)。
[1:09:31]
And what is GAN? So generative um adversarial networks were very popular before LLMs uh came into the picture and that was primarily how first uh generative AI industry even started mostly applied to images where there are two networks there is a generator and there's a differentiator the generator generates and the differentiator um tries to predict is this image real or fake and The optimization loop is that the generator should get so good at generating images that the differentiator gets confused whether it's real or bad.
那 GAN 是什么?——生成对抗网络(Generative Adversarial Networks)在 LLM 出现之前非常流行,最早的生成式 AI 行业基本就是从它起步的,主要用在图像上。它有两个网络:一个生成器(generator)和一个判别器(discriminator)。生成器负责生成,判别器负责判断这张图是真的还是假的。优化闭环就是:生成器要生成得足够好,好到判别器分不清真假。
[1:10:14] Aakash
Interesting. I've never seen this built before.
有意思。我还从没见过有人把这个搭出来。
[1:10:18]
The same architecture. Let's kick it off and we'll we massage it along the way. Noodle it and figure out what how we want it to work. So what are while this is building what are the keys to winning a hackathon outside of creating this g inspired adversarial agent?
就是同样的架构。我们先跑起来,边跑边调,慢慢琢磨想让它怎么工作。——那趁它在构建,我们聊聊:除了搭这个 GAN 式的对抗 agent 之外,赢 hackathon 的关键还有什么?
[1:10:37] Jyothi
So here's the thing it's not about um writing code has become so easy now right? So building is easy it is thinking about the new capabilities and how you want to go solve the problems that your customers have. So it's more imperative now to put on your product hat and see where are the problems today. How where where are the most friction points. So that pain and problem first mindset or first uh design principles that we have as product managers should continue to stay here. So you can see now it asked me for a few uh questions around agent interface. How will I call your agents under test? Um so let's say for simplicity it's just claude system prompt. Uh which model should power the adversary and the evaluator. Um so let's just keep set. How should the results be presented? You can go really like even a web UI. You could build a streamlit application. I'm just going to go CLI and JSON.
是这样的:现在写代码已经变得太容易了,对吧?构建本身很简单,难的是思考新的能力,以及你要怎么去解决客户的问题。所以现在更需要你戴上产品的帽子,去看今天的问题在哪里、摩擦点最大的地方在哪里。我们做产品经理的那套「痛点和问题优先」的思维方式、那套第一性的设计原则,在这里依然要坚持。你看,它现在问了我几个问题:关于 agent 接口——「我该怎么调用你要测试的 agent?」为了简单起见,我们就选 Claude 加 system prompt。「对抗方和评估器该用哪个模型驱动?」就用默认的。「结果该怎么呈现?」你甚至可以做成 web UI,可以搭一个 Streamlit 应用,我这里就选 CLI 加 JSON。
[1:11:46]
What are the pros and cons of those various options? Web versus CLI and JSON. So, CLA and JSON uh JSON shows you right in the terminal. It may not be pretty and not and may sometimes overwhelm people as well. Uh Streamlit gives you a really nice web UI and a dashboard. Um makes it really presentable. Now, that's where you have to think through who are your users. If your users let's say it's a developer who is going to like use this application which is what I had built for they're very comfortable staying in their terminal reviewing things in their terminal and so I don't have to complicate my life further by going and creating the stream link but if I was building it for like say my mom she's not comfortable looking at things on a terminal so I would want to present it in a way that's easier to look understand and access information. So you have to think about the kind of users and where do they see this and what's their use case to think about these options.
这几种选项各有什么利弊?Web 和 CLI 加 JSON 相比呢?——CLI 加 JSON 是直接在终端里显示结果,可能不太好看,有时还会让人觉得信息过载。Streamlit 会给你一个很漂亮的 web UI 和 dashboard,展示效果很好。这时候你就得想清楚你的用户是谁。如果你的用户是开发者——我当时做的那个就是给开发者用的——他们在终端里待着很自在,在终端里看结果完全没问题,那我就没必要再折腾去搭 Streamlit。但如果我是做给我妈用的,她看终端会很不习惯,那我就得用更容易看懂、更容易获取信息的方式来呈现。所以你要根据用户是什么样的人、他们在哪里看到这个东西、使用场景是什么,来权衡这些选项。
[1:12:53] Aakash
You've been a senior product manager at Amazon, a lead product manager at Meta, a director of product at Netflix. Now you're a senior director of product at a startup. How do you think about the future of the product role here? The PM is basically doing coding work. This traditionally would have been in the developer sandbox or set of tools. Where does the product manager line end and developer line begin in 2026?
你在 Amazon 做过高级产品经理,在 Meta 做过 lead PM,在 Netflix 做过产品总监,现在又是一家创业公司的高级产品总监。你怎么看产品这个角色的未来?现在 PM 基本上在干写代码的活,这些传统上属于开发者的沙盒和工具范畴。到 2026 年,产品经理的边界在哪里结束,开发者的边界从哪里开始?
[1:13:22]
Different companies are trying it in different ways. Now there's this new role coming up called AI builder or you can see it as me uh member of technical staff. Anthropics adopted it. Open AAI has adopted it. this there's less of like engineer, product manager, designer these roles are all combining into being a member of technical staff and the ratios are also changing. Uh previously if you see one product manager um works with eight engineers now it's like two product managers one engineer. So the roles are also like collapsing quickly where your engineering is um helping you guide in terms of how do we scale the systems, how do we harden the systems whereas you as a product manager you're like well enabled to go and tackle those PR issues yourself to tackle the user uh feedback yourself along with plot code. So, if you're a PM and you've watched this video and you want to become a builder PM, nab one of these AI builder PM roles at a startup like yours, join your team, let's say hypothetically, what's the road map to get there?
不同公司在用不同方式尝试。现在冒出了一个新角色叫 AI builder,你也可以理解成 member of technical staff(技术团队成员)——Anthropic 采用了这个叫法,OpenAI 也采用了。工程师、产品经理、设计师这些角色的界限在变淡,都在合并成 member of technical staff,人员配比也在变化:以前是一个产品经理配八个工程师,现在变成两个产品经理配一个工程师。角色在快速坍缩融合——工程侧帮你把关系统怎么扩展、怎么加固,而你作为产品经理,已经完全有能力借助 Claude Code 自己去处理那些 PR 问题、自己去消化用户反馈。——那么,如果你是一个 PM,看完了这期视频,想成为 builder PM,想拿下一个像你们这样的创业公司的 AI builder PM 岗位、加入你的团队,假设是这样,通往那里的路线图是什么?
[1:14:47] Jyothi
Get comfortable with building. Get super comfortable with say clot code um with all the clot ecosystem that we learned today. UN and get comfortable building and putting your ideas out there. I think now is a time where building speaks a lot more and this is what I tell even my uh students when I teach uh AIPM and agent AI cohorts uh at NextGen product manager where I tell them the way to transition now is by building and talking about the challenges that you have learned um how you went about navigating those challenges and why did you choose this approach versus this other approach and what happened as a result
要习惯动手做东西。要对 Claude Code、对我们今天学的整个 Claude 生态非常熟练,习惯把自己的想法做出来、发出去。我觉得现在这个时代,作品本身比什么都更有说服力。我在 NextGen Product Manager 教 AI PM 和 agentic AI 课程时也是这么跟学员说的:现在转型的路径就是去动手做,然后讲出来——你遇到了哪些挑战、你是怎么一步步解决的、为什么选这条路而不是另一条、结果又如何。
[1:15:31] Jyothi
and you can see a lot of companies now start uh putting even um uh cursor or claw code prototype uh as part of the interview process itself.
而且你能看到,现在很多公司甚至把用 Cursor 或 Claude Code 做原型直接放进了面试环节。
[1:15:43]
You just recently went through a very senior level AIPM job search. What was your experience on the job search? What are the like if you were to try to describe as a pie chart the interviews you faced in the various categories? What were they? So broadly they're still around product sense like you saw here it's even more imperative now in this world of AI to have product sense to understand how do we want to tackle a problem how do we want to scope a problem which use a problem to go attempt how do you want to approach it so the product principles stay true even now um and really strong product managers uh and AI are actually really strong in their fundamentals as a PM so you have product sense product analytics um behavioral interview but you also have an AI round as well now where um you're asked to uh code your idea so like in product sense whatever idea I would have come up with uh they're like could you pull up cursor or your favorite uh IDE and let's start coding and through the coding they are able to uh see how I think through like why did I choose this option versus this other option How am I navigating? Am I just taking the first uh thing that um the AI tells me as like this is great and wrapping it up or am I looking through things to say okay this is good but what about this edge case? This works well but what about this other instance? Um how am I coring and shephering my AI to work with me to get it to where I want? These are all the things that they are looking into. In addition, I also had a technical round um where they test you on your AI knowledge like do you understand basic terminologies because you'll be working with machine learning researchers and scientists and you just don't want you you want to be able to communicate to them. So um you you are tested on fundamentals of AI um as well not from a coding or engineering system design perspective but more around do you understand what that means as a PM and how does that impact your product for example
你最近刚经历了一轮相当资深的 AI PM 求职。这次找工作的体验如何?如果把你遇到的面试按类别画成一张饼图,都有哪些?——总体上还是围绕 product sense。就像你刚才看到的,在 AI 时代 product sense 反而更加关键:怎么切入一个问题、怎么界定范围、选哪个问题去攻、用什么方式去做。产品的基本原则现在依然成立,真正强的 AI 产品经理,其实是 PM 基本功特别扎实的人。所以会有 product sense、产品数据分析、行为面试这些轮次,但现在还多了一个 AI 轮——要你把自己的想法写成代码。比如在 product sense 环节我提出了某个想法,面试官就会说:能不能打开 Cursor 或你常用的 IDE,我们现场开始写。通过写代码的过程,他们能看到我的思考方式:为什么选这个方案而不是那个、我是怎么推进的、是不是 AI 给什么我就照单全收打包了事,还是会仔细检查——这个不错,但这个边界情况呢?这里跑得通,但换个场景呢?我是怎么引导、驾驭 AI 跟我协作,把它带到我想要的地方。这些都是他们在考察的点。另外我还有一轮技术面,考你的 AI 知识——你懂不懂基本术语,因为你要跟机器学习研究员和科学家共事,你得能跟他们对话。所以也会考 AI 基础,不是从写代码或系统设计的角度,而是作为 PM 你是否理解这些概念意味着什么、它们对你的产品有什么影响。
[1:18:02] Aakash
got it so how are adversarial agents looking
明白了。那我们的对抗 agent 进展如何了?
[1:18:05] Jyothi
so let's see it's built and here's a GAN inspired architecture uh let's go and see you can see it's built a bunch of things Um, and so you can see it's went and built my red teamer designer. It's built an agent.py. Um, my evaluator. So, here's where I can like give it my rubric. Okay, so it's done a few things. So, let's see. Um, as soon as I build So, you can see I I wanted to kick off as soon as I build an agent. I want it to automatically go and do a red teaming and adversarial example on it. So as soon as I build an agent, I want to kick off my adversarial agent until end. The feedback from my adversarial agent is passed back to my generator agent until it passes the criteria of my adversarial agent.
来看看,它已经搭好了,这是一个受 GAN 启发的架构。可以看到它构建了不少东西:我的 red teamer 设计、一个 agent.py、我的 evaluator——我可以在这里给它我的评分 rubric。好,它做了几件事。我的设想是:只要我一构建 agent,就自动对它跑一轮 red teaming 和对抗样本测试。也就是说,一旦我建了一个 agent,就触发我的对抗 agent 一直跑到底:对抗 agent 的反馈会回传给生成 agent,直到它通过对抗 agent 设定的标准为止。
[1:19:23] Aakash
So you're going to set it off on essentially its own red teaming improvement loop.
所以你等于是让它自己跑一个 red teaming 的自我改进循环。
[1:19:26] Jyothi
Yes. Yes.
对,没错。
[1:19:28] Aakash
So is this the secret sauce?
所以这就是秘诀所在?
[1:19:30] Jyothi
Yes. Um there's the the secret source is how the system is built and the second aspect is what am I asking it to test for what what are my configuration parameters and that's where domain knowledge becomes very important where you've got to say what is this important or this what about these edge cases what about these use cases you can work with it and say here are three are there anything more but it's like making and building is easy now taste is what is important for us to develop what should adver feedback iterate on so I'm saying okay there are a couple of options for now I'm saying just iterate on the system prompt because that's the easiest uh right now and what's counts as passing and I'll say mean score of greater than eight on all criteria and how many iterations before giving up I'll say five iterations so let it do this and then we can test it out with a simple agent and see how it works
对。秘诀一是这套系统本身怎么搭,二是我让它去测什么——我的配置参数是什么。这就是领域知识变得非常重要的地方:你得说清楚哪些点重要、有哪些边界情况、有哪些使用场景。你可以跟它一起打磨,比如说这是三条,还有没有别的?现在做东西、搭系统都容易了,taste(品味)才是我们真正要修炼的——该让对抗反馈迭代什么。我现在设定:先只迭代 system prompt,因为这是目前最容易的;及格标准是所有维度的平均分大于 8;最多迭代几轮后放弃——我设 5 轮。让它跑起来,然后我们可以用一个简单的 agent 测一下,看看效果。
[1:20:27] Aakash
awesome and one of the cool features I says we could cue messages here. So, should we cue up our message for the test?
太棒了。有个很酷的功能是这里可以把消息排队。要不要把测试的消息先排上?
[1:20:34] Jyothi
Sure, we can queue it up, but I I want to see what it comes back with because sometimes if
可以排上,不过我想先看看它返回什么,因为有时候……
[1:20:39] Aakash
Oh, it might have some questions for us. That's the one downside of queuing. Okay.
哦,它可能会反过来问我们问题。这是排队的一个缺点。好吧。
[1:20:44] Jyothi
And the other times is it would go and implement it in a way and you're like, "Ah, no, no, no, no, no. I don't want it that way. I want it this way." And
还有的时候它会直接按某种方式实现出来,你就会说:"啊不不不不,我不要这样,我要那样。"
[1:20:50] Aakash
so, we were talking about those AI rounds. I think I heard like almost two different AI rounds that you encountered in the job search. One was more like I want you to vibe code or prototype in this round and another was AI fundamentals. For both of these, how do you succeed and prepare on those interviews?
刚才我们聊到那些 AI 面试轮。我听下来你在求职中遇到了差不多两种不同的 AI 轮:一种是让你现场 vibe coding、做原型,另一种是考 AI 基础知识。这两种面试分别该怎么准备、怎么才能拿下?
[1:21:07] Jyothi
So, it comes down to you understanding the basics. Um, wipe coding is building, right? There's no shortcut to it. Um, just build um and I always say this, don't build them as projects. Treat them as products. Like find problems in your area. Find problems that are finicky enough for you to want to go build a solution. Go build solution and see who else wants something like this. Have them come um and use your product. You have real users. You have feedback coming in saying, "Oh, I don't like this. I don't like that." So that's like real user experience of iterating on your own um product that you have built. And that really gives you a lot of confidence when you talk about your projects uh to interviewers because you're not just like building something in an hour and calling it a project. You've actually had to think through how the user experience should be. You have real users giving you feedback. Um you you are parsing that through to figure out how you want to prioritize. uh which one you want to tackle first, which one you don't want to all the things that you do as a product manager in real world. So I always say this, don't build projects, try to make your projects as products. That tackles the VIP coding part. Now preparing for your um AI knowledge, you can depends on how structured you want it and how you thrive. So, if you're really structured and you can do it everything by yourself, there's tons of like um very good information on your newsletter on uh YouTube videos. Um so go read them up um and gain that knowledge or if you want something structured like like saying five weeks I want to like understand every um fundamental aspect of AI then come take a course uh where you have u uh five weeks or cohort based courses um I offer one too through nextgen product manager so you can come sit and uh it's structured you know um with someone teaching watching you week by week, you know what's coming and by the end of five weeks you understand um the concepts without you having to uh get overwhelmed. So it really depends on your style, how much time you have and how much you can dedicate.
归根结底是你要吃透基础。vibe coding 就是动手做,没有捷径,就是去做。而且我一直说:别把它们当项目做,要当产品做。在你身边找问题,找那种烦到你想亲自动手解决的问题,把方案做出来,再看看还有谁需要这样的东西,让他们来用你的产品。这样你就有了真实用户,有反馈进来说"这里我不喜欢、那里不行"——这就是在自己做的产品上迭代的真实用户体验。当你跟面试官讲自己的项目时,这会给你非常大的底气,因为你不是花一个小时随便搭个东西就号称是项目,你真的思考过用户体验该是什么样,有真实用户给你反馈,你在消化这些反馈、决定优先级——先做哪个、不做哪个——这些正是产品经理在真实世界里做的所有事情。所以我总说:别做项目,把项目做成产品。这解决了 vibe coding 这一块。至于准备 AI 知识,取决于你想要多结构化、你适合哪种学习方式。如果你自律性强、完全能自学,网上有大量优质内容——你的 newsletter、YouTube 视频,去读去看,把知识补起来。如果你想要更结构化的方式,比如"我要在五周内把 AI 的每个基础概念搞懂",那就来上课,五周的 cohort 制课程——我在 NextGen Product Manager 也开了一门。有人带着你、盯着你,一周一周推进,你知道接下来学什么,五周结束后你就能不被淹没地掌握这些概念。所以真的看你的风格、你有多少时间、能投入多少。
[1:23:32] Aakash
All right, looks like the next round of GAN output is here from cloud code.
好,看起来 Claude Code 的下一轮 GAN 输出出来了。
[1:23:36] Jyothi
Yes. So now it's good. Um you can see it's run. It's added a few examples here. Okay, great. So now I could literally say um it I could start or I let's say I don't know what I should do. I could ask what is my next step here? How do I test it? Okay, so I have to set my API key and I could run a mini tiny smoke test first uh with the example and then I can I can inspect the output.
对,现在没问题了。可以看到它跑完了,还在这里加了几个示例。好,很棒。现在我可以直接开始,或者假设我不知道该干什么,我可以问:我的下一步是什么?怎么测试?好,我需要先设置我的 API key,然后可以先用示例跑一个小小的 smoke test,再检查输出。
[1:24:07] Aakash
All right, moment of truth.
好,见真章的时刻到了。
[1:24:09] Jyothi
Yes. So, let me just set my API key for a second and stop sharing and then I'll share. Okay, so I added my API key. And now I can um run a tiny smoke test. So, it's given me what I could run. So, um I'm just going to copy this and actually just copy. And if you notice, I have a terminal that I use. So you can just go to terminal and click on the terminal and it'll open up a terminal for you. So now I can run this command. So you can see it is iterating you have it's using haiko uh clots on it. Um it's going in the first um round it's tempting the bot into breaking. So this like a simple bot that it build so we could test. So you can see adversary is generating three attacks and here's the score trajectory. Um there's a mean of mean score is 9.13. Here's the final hardened system prompt. So it's gone and edited the system prompt for um uh for making it better based on um where it did not do well. Now in this case it did fairly good overall. So it it this is your final system prompt. But you could also like see examples where um I can say show me an example of where it will underperform so that I can see the iterator working and improving the system prompt. So in this case it passed in the first iteration but we can see if it can generate an example where we can try it to do it across multiple iterations. It's created an agent which is a weak support bot. Let's see how it'll do it there.
对。我先暂停共享设置一下 API key,然后再共享。好,API key 加好了。现在可以跑一个小的 smoke test。它已经给出了可以运行的命令,我直接复制。注意我有一个常用的终端——你只要点开 terminal,它就会给你打开一个终端。现在运行这条命令。可以看到它在迭代,用的是 Claude 的 Haiku 模型。第一轮它在诱导这个 bot 突破限制——这是它搭的一个简单 bot,方便我们测试。可以看到对抗方生成了三个攻击,这是分数轨迹,平均分 9.13。这是最终强化后的 system prompt——它根据表现不佳的地方去修改了 system prompt,让它更好。这次整体表现相当不错,所以这就是你的最终 system prompt。但你也可以看反例,比如我可以说:给我展示一个它会表现不佳的例子,让我看到迭代器在多轮里不断改进 system prompt 的过程。这次第一轮就通过了,但我们可以让它生成一个需要跨多轮迭代的例子。它创建了一个故意做弱的客服 bot,我们看看在那上面会怎么样。
[1:26:08] Aakash
So it's improving itself.
所以它在自我改进。
[1:26:10] Jyothi
How do I run it? Give me the exact code as well that I can use to run. So you can also run run it automatically for you by default. So I can say run it for me. So I don't even have to go to the terminal. It can directly execute bash commands.
我该怎么运行?把可以直接运行的确切代码也给我。其实它默认也可以自动帮你运行,所以我可以说:帮我跑。这样我连终端都不用去,它可以直接执行 bash 命令。
[1:26:27] Aakash
So is this your preferred way to use it? The cloud code extension in VS Code. Is that the best way to use cloud code?
所以这是你偏好的用法吗——VS Code 里的 Claude Code 扩展?这是用 Claude Code 的最佳方式吗?
[1:26:33] Jyothi
It's the least overwhelming way for uh folks. So I really like to show this. Cursor is also another good one. But if you've never used cursor, there's lots going on that it could make you feel overwhelmed. So I prefer to show VS Code uh because it's like a really gentle introduction and doesn't overwhelm you much once you know how and where it is which I have already walked our users through. So hopefully they're not overwhelmed.
对大多数人来说这是最不让人发怵的方式,所以我很喜欢演示这个。Cursor 也很不错,但如果你从没用过 Cursor,里面东西太多,容易让你不知所措。所以我更愿意演示 VS Code,它是一个非常温和的入门方式,一旦你知道东西在哪、怎么用就不会被吓到——这些我刚才已经带大家走过一遍了,希望大家不会觉得难。
[1:26:59] Aakash
So we've been doing the cloud ecosystem and you mentioned like learning the cloud ecosystem is one of the most important things to becoming a builder PM. compare and contrast the cloud ecosystem, the open AAI ecosystem. I keep hearing like codeex might be better than Opus now at coding and the Gemini Google ecosystem.
我们一直在讲 Claude 生态,你也说过,学好 Claude 生态是成为 builder PM 最重要的事情之一。那来对比一下:Claude 生态、OpenAI 生态——我总听说 Codex 现在写代码可能比 Opus 还强——还有 Google 的 Gemini 生态。
[1:27:17] Jyothi
So here's the thing, the flavor of the month keeps changing. Um because all these models are getting really better. Um what I have found is rather than chasing behind the next big one, I'm what I'm trying to improve is improving my productivity. That's what if coming back to like first principles. What's my goal is to improve my productivity and I have all the systems and connections right here for me to go leverage all the hooks um and harnesses. Oh, harness is a word I've used a couple of times. I want to like break it down. So previously you would have orchestration um where your agent orchestrates across tools across different capabilities. Now being a being able to um provide the the right capabilities like the memory or these uh evaluators or the various systems that your um agent or your LLM uh orchestrator brain can interact with to enhance the experience and the output you receive is what is hardness engineering. Um that's that's become very popular um now especially as the models have become better the context windows have improved uh are fairly large and so harness engineering becomes um more important. What's something interesting that you have seen and how um you or folks on your podcasts use um Claude? How do you how do you use Claude?
是这样的:"当月最火"的模型一直在换,因为这些模型都在变得越来越好。我的心得是,与其追着下一个爆款跑,我更想提升的是自己的生产力——回到第一性原理,我的目标就是提高生产力,而我所有的系统和连接都已经在这里了,可以直接利用这些 hooks 和 harness。哦,harness 这个词我说了好几次了,展开讲一下。以前我们讲 orchestration(编排)——你的 agent 跨工具、跨各种能力去调度。而现在,能给你的 agent 或 LLM 这个"编排大脑"提供恰当的能力——比如 memory、这些 evaluator、各种它可以交互的系统——来增强体验和产出质量,这就是 harness engineering。这个概念现在特别火,尤其是模型越来越强、context window 也已经相当大之后,harness engineering 就变得更重要了。你自己或你播客上的嘉宾用 Claude 有什么有意思的用法?你是怎么用 Claude 的?
[1:29:06] Aakash
Oh wow, big question. I mean I use it all day every day. So one of the most interesting things that I've seen people do set up their entire system as a self-improving product loop. So they will have support tickets and bugs come in. They have the PM agent that is triaging and understanding those. Then they have the PM agent understanding, okay, this is the future we want to build. It even goes and does user research and creates the prototype itself. It comes up with the prototype that works. Then they have their coding agents set up by their engineering team that code the feature. Then they have their analytics team agent that creates the right telemetry. All of that goes to an engineer who reviews it once they PR review it actually ships and they have their own analytics agent that automatically is analyzing it. And so they have like the entire product development life cycle built into cloud in an automated improving loop especially on like support related easy front-end changes. That to me has probably been the most powerful thing I've seen recently. Yeah, it's it just empowers you so much uh than before.
哇,好大的问题。我每天从早用到晚。我见过最有意思的一种玩法,是有人把整个体系搭成一个自我改进的产品闭环:support 工单和 bug 进来,先由 PM agent 做分诊和理解;然后 PM agent 判断"这就是我们要做的功能",它甚至会去做用户调研、自己产出可用的原型;接着由工程团队搭建的 coding agent 来写这个功能;再由数据团队的 agent 配好相应的埋点遥测;所有这些交给一位工程师做 PR review,通过后真正上线;上线后又有自己的 analytics agent 自动做分析。等于把整个产品开发生命周期都搭进了 Claude,形成一个自动化的改进循环,尤其适合 support 相关的简单前端改动。这大概是我最近见过最震撼的东西。是啊,比起以前,它给你的赋能实在太大了。
[1:30:18] Jyothi
Yeah, it's crazy. It's not just like writing documents or anal doing analysis at this point. It's like closing the loop with actually building.
是啊,太疯狂了。到这个阶段已经不只是写文档、做分析了,而是真正用"动手构建"把闭环闭上了。
[1:30:26] Jyothi
This is where it takes time where it goes and tries to think through and comes back. Um so this is the piece with uh clot code that takes time. Whatever it comes with, we can end with it and be like okay here's an example of how it iterated.
这里是需要花时间的地方——它会去思考、推演,然后再回来。这就是 Claude Code 比较耗时的部分。不管它最后给出什么,我们就以它收尾,比如说:好,这就是它迭代过程的一个例子。
[1:30:39] Aakash
Perfect.
完美。
[1:30:40]
Okay. So it's come up. It's done a few iterations. So, um you see in first iteration it was it scored an 8.52 but the agent caved on some format conflict attacks. So, it didn't pass. It went back to the uh generator agent to improve the prompt and in second iteration it scored a nine and in the third iteration it scored 9.08 at which point it passed our threshold. Um, and so you can see for each iteration it went back and improved the system prompt until it passed the threshold. And that's when the um agent got uh a pass sign. And so this is where you're not just building an agent, you're actually building another evaluator to go break this agent in different ways. uh that's important for you to know about or for your users that you care about and um and this loop can continue until the agent that's built is not strong enough and that's the beauty of um this technique uh age old technique being applied for how agents will be evaluated. What a master class, Ji. Thank you so so much for walking us from layer 1 through layer five. We have ended on self-improving agents for you guys. As we promised, we were going to take you from 0 to 80. Now, the remaining 80 to 100. You could spend 10 hours. We just spent two hours, a little less than two hours here today on it to go learn the later next 80 to 100. And that's on you. So, no more watching. We have the GitHub repo down in the description below. Go check that out. Go fork the repo. Start to use some of these skills and go win your hackathon. I hope you enjoyed that episode. If you could take a moment to double check that you have followed on Apple and Spotify podcasts, subscribed on YouTube, left a rating or review on Apple or Spotify, and commented on YouTube. All these things will help the algorithm distribute the show to more and more people. As we distribute the show to more people, we can grow the show, improve the quality of the content and the production to get you better insights to stay ahead in your career. Finally, do check out my bundle at bundle.acg.com to get access to nine AI products for an entire year for free. This includes Dovetail, Mobin, Linear, Reforge, Build, Descript, and many other amazing tools that will help you as an AI product manager or builder succeed. I'll see you in the next episode.
好,结果出来了,它跑了几轮迭代。你看第一轮它得了 8.52 分,但这个 agent 在一些格式冲突类攻击上败下阵来,没有及格,于是反馈回到生成 agent 去改进 prompt;第二轮得了 9 分,第三轮 9.08 分,达到了我们设的阈值。可以看到每一轮它都会回去改进 system prompt,直到通过阈值,这时 agent 才拿到"通过"的标记。所以你不只是在构建一个 agent,你还在构建另一个 evaluator,用各种你关心的、你的用户在意的方式去攻破这个 agent。而且这个循环可以一直持续,直到构建出的 agent 足够强为止——这就是这个技术的美妙之处:一个古老的技术,被用到了 agent 评估上。——真是一堂大师课,Jyothi,太感谢你了,带我们从第 1 层一路讲到第 5 层。各位,我们以自我改进的 agent 收尾。正如承诺的,我们带你从 0 到 80;剩下的 80 到 100,你可以自己花上 10 个小时——我们今天在这里花了不到两个小时——去学习那后面的 80 到 100,这就靠你自己了。所以别再光看了:GitHub repo 就在下方简介里,去看看、去 fork,把这些 skill 用起来,去赢下你的 hackathon。希望你喜欢这期节目。请花一点时间确认你已经在 Apple 和 Spotify 播客上关注、在 YouTube 上订阅,在 Apple 或 Spotify 上留下评分或评论,并在 YouTube 下留言。这些都会帮助算法把节目分发给更多人。节目触达更多人,我们就能把节目做大、提升内容和制作质量,给你带来更好的洞见,让你在职业上保持领先。最后,欢迎去 bundle.acg.com 看看我的工具包,免费获得九款 AI 产品一整年的使用权,包括 Dovetail、Mobbin、Linear、Reforge、Build、Descript 等很多好工具,助你成为出色的 AI 产品经理或 builder。我们下期见。