The AI PM Skill That Gets You Instant Job Offers | Aparna, Arize AI
频道: Aakash Gupta
视频: https://www.youtube.com/watch?v=DL-pUGcfrf4
原文语言: en
统计: 共 51 轮 · Aakash 12 · Aparna 38
[0:00] Aakash
Any product person that has [music] used observability and is looking at their traces and looking at their evals, you're probably already in the top 1% of PMs. What is the role then of the PM? Like, do PMs need to become engineers at this point?
任何一个用过 observability、会去看自己 traces、会去看 evals 的产品人,你大概率已经是 PM 里的前 1% 了。那 PM 这个角色到底还剩下什么?比如说,PM 到现在是不是得变成工程师了?
[0:13] Aakash
AI native teams, [music] I am seeing that the gap between a PM and an engineer is indistinguishable. Aparna Dhinakaran is the CPO [music] and co-founder of Arize AI. $131 million raised and most of the smartest AI teams [music] I know building their evals on top of it.
在 AI native 的团队里,我看到 PM 和工程师之间的界线已经几乎分不出来了。Aparna Dhinakaran 是 Arize AI 的 CPO 兼联合创始人。公司融了 1.31 亿美元,我认识的最聪明的那批 AI 团队,evals 几乎都搭在它上面。
[0:31] Aakash
like a good eval is like you're getting some healthy percentage right, but also healthy wrong so that you can make progress, right? [music]
好的 eval 是这样的:它有相当一部分判对,但也有相当一部分判错,这样你才有空间去往前推进,对吧?
[0:37] Aparna
100%. Like, I get excited when I see that evals are wrong because then it gives me a chance to know that there's [music] improvement that could be made.
百分之百同意。我看到 evals 出错的时候反而会很兴奋,因为这正好告诉我:这里有可以改进的地方。
[0:45] Aakash
What are the things if somebody has just 2 hours this weekend that they should concretely go do and take away besides just they've [music] watched this episode, but now they're going to actually make impact in their career?
如果有人这个周末只有 2 个小时,除了看完这一期节目之外,他们具体可以去做点什么、带走点什么,真正对自己的职业生涯产生影响?
[0:55] Aparna
If you have any 2 hours this weekend, [music] I would say literally what we just did right now, which is Before we get into today's episode, I wanted to share that you can get a free year of my favorite AI tools including Bolt that new, Mobbin, Arize, Relay app, Dovetail, Linear, Magic Patterns, Reforge Build, Descript, and Speechify if you join my bundle at bundle.akashg.com. On top of that, I wanted to quickly ask you to please double-check that you are subscribed on YouTube, Apple, and Spotify podcasts. It's a free thing you can do that really helps support the show. And now into today's episode. So, I've been doing a ton of episodes on Claude code, a ton of episodes on AI agents, and separately episodes on evals. What this episode we're doing today is we're bringing it all together for you in one iterative loop. It's kind of like the product development cycle for AI products in a single shot. So, you're going to get to see front to back how we do it. I think we have a tremendous opportunity to learn from Aparna. So, I'm going to try to ask her the tough questions for you guys where maybe what she's doing she's skipping some steps so that you guys can see it step by step. And she's volunteered to be our guinea pig on this. So, Aparna, thank you so so much for showing us the ropes of how to do Claude coding evals. I'm super super excited to be here. Thanks so much for having me, Akash.
如果这个周末你有 2 个小时,我会说,就照着我们刚刚做的这一套来做。在进入今天的正片之前,我想先告诉大家:如果你加入我的 bundle(bundle.akashg.com),可以免费用一年我最喜欢的那些 AI 工具,包括 Bolt(那个新的)、Mobbin、Arize、Relay app、Dovetail、Linear、Magic Patterns、Reforge Build、Descript 还有 Speechify。另外,也顺便拜托大家在 YouTube、Apple 和 Spotify 播客上确认一下有没有订阅——这是个免费的小动作,但对节目帮助真的很大。好,进入今天的正片。我之前做了一大堆关于 Claude Code 的节目、一大堆关于 AI agent 的节目,还单独做过几期讲 evals 的。今天这一期,我们要把这些全部串成一个完整的迭代闭环给你看,差不多就是把 AI 产品的整个产品开发周期一口气走完。所以你能从头到尾看到我们是怎么做的。我觉得我们有一个特别好的机会,可以从 Aparna 身上学到东西。我会替你们去问那些比较扎实的问题——可能有些步骤她做着做着会跳过去,我会让她拆开、一步一步讲清楚。她自己主动当我们的小白鼠。所以 Aparna,太感谢你来带我们走一遍怎么用 Claude 做 coding 和 evals。我超级开心能来,非常感谢你邀请我,Aakash。
[2:30] Aakash
So, what are people getting wrong when you look at them building Claude code agents and trying to do evals? Yeah, I mean, I think the first question I get asked a lot is when should I even start doing evals? Like, why is that important? Um do you do I need to think about it before I even build my agent? And I mean, if I'm honest with you, most teams are starting a a you know, they're starting with just building. Like, you got to start by having a real product before you want to you know, you you run evals on it. And so, um today what I'm going to actually walk you through is the full end-to-end loop of getting started with building a product, when does it make sense to actually, because of the data that you've collected, start to actually run evals and automate that. Awesome. Let's see it in action. Where should we start? So, it's a little bit of a vision for anyone who's an AI PM today. Code is so cheap to go create, which means that product taste is really the alpha today. People, especially product managers, there's all this hype around, you know, are is it going to be the death of PMs? You know, I'll tell you this, we're hiring more PMs than ever. We're hiring more engineers than ever. The ones that stand out are those that actually have an opinion and a taste around what to go build. And so, today, you know, a little cheeky, but um can we try to create taste? Can we try to have the PMs that are watching this have a upper hand to actually create that product taste? Well, where do where does product taste actually come from? You know, you look at kind of some of the best products out there. And what they're doing is taking in a ton of feedback. I mean, the best PMs do this. Best PMs, I mean, YC says this uh to their eight to every single
那你观察下来,大家在用 Claude Code 搭 agent、还想做 evals 的时候,通常会犯什么错?嗯,我觉得我最常被问到的第一个问题是:我到底该从什么时候开始做 evals?为什么这件事重要?我是不是在还没搭 agent 之前就该想这个?老实跟你说,大多数团队其实就是先动手搭——你得先有个真正的产品,然后才谈得上在上面跑 evals。所以今天我要带你走的是一整条端到端的闭环:先从搭产品开始,然后看在什么时候——因为你已经攒了足够的数据——开始真正去跑 evals、把它自动化起来。太棒了,咱们实操看看。从哪儿开始?这里先给今天所有 AI PM 一点大方向。代码现在生成成本极低,这意味着今天真正的 alpha 是产品品味(product taste)。大家,尤其是产品经理,现在到处都在炒作那句话:是不是 PM 要消亡了?我跟你说,我们现在招的 PM 比以往任何时候都多,招的工程师也比以往任何时候都多。真正脱颖而出的,是那些对该做什么有自己的观点、有自己品味的人。所以今天,说得有点俏皮,咱们能不能试着把品味造出来?能不能让正在看这期的 PM 们,拿到一个真正去打造产品品味的先手优势?那产品品味到底从哪儿来?你去看市面上最好的那些产品,它们在做的事其实是吸收海量的反馈。最好的 PM 都这么干。最好的 PM——YC 对他们每一届
[4:20] Aparna
cohort, which is talk to users and go build. And I think what we see is that in order to actually create taste, you need to be getting feedback from a ton of different sources. From It could be everything from where your team stores those issues. It could be from GitHub discussions or like, you know, in real life discussions. From Slack and Discord, um from your actual community talking to you. But also, we see teams building out really a context graph with all of this feedback. Everything from Gong transcripts, every time you talk to your customers. Your product analytic tools from Posthog and Amplitude and Pendo and FullStory. Um even down to Twitter. If you have a product that your users are tweeting about and sharing feedback on, these are all ways for you to actually create and cultivate that feedback source. And instead of having just a human consume it, you can actually have your agent consume that feedback. And so, what we're going to do today is we're going to build a bit of a product taste agent. This agent, you you're PM, uh your your job is to come in and kind of figure out what to go build. What are users asking for? Every day, this product taste agent's going to tell you what your biggest pains are, what your biggest priority should be, and suggest where your product road map needs to go. Um the product I'm going to work off of today, and you can pick your own product that makes sense for you, but the product I'm going to pick is actually our own open source product, Arise Phoenix. Arise Phoenix is the leading open source observability and Evals platform. You can actually get started and host everything entirely open source with Phoenix. Um, but with Phoenix, and you're going to see what I do here, is that we have a
每一届都会说同一句话:去跟用户聊,然后去搭东西。我觉得我们看到的是,要真正培养出品味,你得从一大堆不同的来源去收集反馈。可以是你团队存放问题的地方,可以是 GitHub discussions,或者就是现实里的讨论。从 Slack 和 Discord,从你那些真正会找你聊的社区。但我们也看到有的团队会把这些反馈搭成一张 context graph(上下文图谱):从每次跟客户聊完的 Gong 转录,到你的产品分析工具——Posthog、Amplitude、Pendo、FullStory,甚至一直到 Twitter。如果你的产品有用户在上面发推、分享反馈,这些全都是你去创造、去养出这个反馈来源的渠道。而且,与其只让一个人去消化这些反馈,你完全可以让你的 agent 去消化。所以今天我们要做的,是搭一个产品品味 agent。这个 agent——你是 PM,你的工作是进来想清楚该做什么、用户在要什么——每天它都会告诉你:你现在最大的痛点是什么、最该优先做的是什么,并给你的产品路线图提建议。今天我要拿来做示范的产品——你完全可以挑一个对你自己有意义的——我要挑的其实是我们自己的开源产品 Arize Phoenix。Arize Phoenix 是目前最主流的开源 observability 和 evals 平台。你完全可以用 Phoenix 把整套东西以开源方式跑起来、自己托管。但用 Phoenix 的时候,你接下来会看到我做的事,是我们有一
[6:09] Aparna
ton of backlog of issues. We also have a really vibrant GitHub discussions. We have our own Slack community. We have feedback from people who are tweeting at us. And so, what I'm going to try to do is actually aggregate a lot of that uh, I'm going to actually try to aggregate that feedback and use that to surface up where should we go and what should we build next. So, the steps we're going to do here is actually first create this PM agent. We're going to do this using Cloud Code. The magic behind everything that we're going to use to improve is really tracing. We're going to trace everything. We're going to get literally every step of what our agent does is going to be visible to us. And then we're actually going to run the Evals after question. And I think this is kind of the big, you know, when people ask, "When do I do Evals?" You know, I always, you know, kind of point towards get the data, trace everything, get the observability. The Evals can kind of help you then take you to the next level for your agent. Um, so, we're going to trace it, we're going to eval it, and then we're going to do this loop where we improve our agent and bring it right back. Um, so, pick your favorite product that you want to actually use. Pick a product that you have all the context of. Um, you could start super simple. What I'm going to start with today is literally just the GitHub, you know, issues, the GitHub discussions, and use that to actually inform what my product taste or PM agent is going to look like. Let's see this. We're going to go ahead and build a PM product taste agent just using Cloud Code. So, go ahead, kick up Cloud code in your terminal. For product folks, you know, this might feel intimidating in the beginning, but I can guarantee you
一大堆积压的 issue。我们还有一个非常活跃的 GitHub discussions,有我们自己的 Slack 社区,还有那些在 Twitter 上 @我们提反馈的人。所以我要做的,是把这些反馈尽量聚合起来,用它来浮现出:我们接下来该往哪儿走、该做什么。这里我们要走的步骤是:第一步,先把这个 PM agent 建出来,用 Claude Code 来建。我们接下来要用来做改进的核心魔法,其实就是 tracing。我们要把所有东西都 trace 下来——agent 做的每一步,对我们都是可见的。然后我们再在这之后去跑 evals。我觉得这正是大家常问的那个大问题:『我到底什么时候该做 evals?』我永远会指向同一个方向:先把数据拿到、把所有东西 trace 下来、把 observability 搭好。Evals 是之后帮你把 agent 带到下一个层次的东西。所以我们会先 trace 它、再 eval 它,然后做这样一个闭环:改进我们的 agent,再把它放回去跑。所以你挑一个你自己最喜欢、又掌握全部上下文的产品。你可以从超简单的开始。我今天就只从 GitHub 的 issue、GitHub discussions 入手,用它们来决定我的产品品味 agent(或者说 PM agent)长什么样。咱们来看看。我们就用 Claude Code 来搭一个 PM 产品品味 agent。来,在你的终端里把 Claude Code 启起来。对产品同学来说,一开始这可能有点吓人,但我可以跟你保证,
[7:52] Aparna
the level of control and iteration you're going to get by just doing this in your terminal and getting comfortable is going to feel just the unlock you're going to get is going to be worth a little bit of that learning kind of pain in the beginning. I'll be honest, I've not always been the best with my email inbox and just thinking about it made me feel anxiety. But my anxiety has really never been lower since I started using Superhuman Mail, today's podcast sponsor. Their Ask AI feature is one thing that really stands out for me because I have so many contract details or deliverables buried eight replies deep and I can just ask the AI. I also love the auto drafts feature so that I have a draft to react and respond to and of course their follow-ups are a lifesaver. Now is the time to give it a try. Check it out at superhuman.com/aakash. Today's episode is brought to you by Vanta. As a founder, you're moving fast toward product market fit, your next round, or your first big enterprise deal. But with AI accelerating how quickly startups build and ship, security expectations are higher earlier than ever. Getting security and compliance right can unlock growth or stall it if you wait too long. With deep integrations and automated workflows built for fast-moving teams, Vanta gets you audit-ready fast and keeps you secure with continuous monitoring as your models, infra, and customers evolve. Fast-growing startups like Link, Chain, Rider, and Cursor trust Vanta to build a scalable foundation from the start. So, go to vanta.com/aakash. That's v a n t a.com/aakash to save $1,000 and join over 10,000 ambitious companies already scaling with Vanta. So, let's do this. Uh go ahead and create a repo or create just a directory and you can go ahead
你光是在终端里这么做、把它玩熟之后,能拿到的那种掌控感和迭代速度——那个解锁,绝对值回你一开始那一点点学习的阵痛。说实话,我以前一直跟邮箱处理得不太好,光是想到它就让我焦虑。但自从我开始用 Superhuman Mail(今天节目的赞助商),我的焦虑感降到了从未有过的低点。它的 Ask AI 功能对我来说特别突出,因为我有太多合同细节、太多交付项埋在第八层回复里,现在我直接问 AI 就行。我也很喜欢它的自动草稿功能,这样我手里总有一份可以拿来回应、修改的草稿,当然还有它的 follow-up 功能,简直是救命。现在正是去试一试的好时候,去 superhuman.com/aakash 看看。今天这期由 Vanta 赞助。作为创始人,你正全速冲向产品市场契合、下一轮融资,或者第一个大企业客户。但随着 AI 让初创公司搭建和上线的速度越来越快,安全方面的要求也比以往任何时候都来得更早。安全与合规这件事做对了能解锁增长,拖太久反而会卡住增长。Vanta 用深度集成和为快节奏团队打造的自动化流程,让你快速达到 audit-ready(可审计)状态,并随着你的模型、基础设施和客户不断演进,用持续监控帮你一直保持安全。像 Link、Chain、Rider、Cursor 这些高速成长的初创公司,从一开始就信任 Vanta 来搭建可扩展的基础。所以去 vanta.com/aakash,也就是 v-a-n-t-a.com/aakash,可以省 1000 美元,加入已经有超过一万家雄心勃勃、正在用 Vanta 扩张的公司。好,咱们开干。先建一个 repo,或者就建一个目录,然后你可以
[9:35] Aparna
and initialize Claude inside of that directory. And let's just go ahead and first give it a starter prompt actually build this agent. I'm going to ask it to build me a PM agent for the Arize AI Phoenix product. Um and I can go ahead and actually just link the URL to that entire repo directly in here. So that it has exactly context of what I'm asking it to build. Um and then I'm just going to go ahead and ask what context do I want it to have? So pull recent GitHub discussions, pull all the recent releases. Um and look at the GitHub issues. I'm going to start kind of piecemeal here first, first just starting with context from one location, which is GitHub. As we scale this, you can add in context from like I was saying, your Gong transcripts, your product analytics, you can add context from um literally your Slack convos, your Discord channels, anything can be brought in here. And what I first wanted to do is first just figure out score the issues and the discussions um based off of priority. Like first just figure out how important is the stuff that we want it to actually look at and build. So things to look at is like bugs versus features, uh reactions that people gave it, comments, you know, I do want it to look at recency. So these are all things that I'm actually asking this product taste agent to take a look at and consider. Uh then call Claude or you know, I can be specific here. I can say call Claude Opus, whatever model I want. So call Claude um with uh you know, I am I could even ask it to go ahead do some kind of like prompt caching so that it doesn't keep pulling down the issues every time that I run this loop, but just to keep it simple in the beginning, what I'm going to do is just call Claude and uh write down a just a markdown PM report that has um
在那个目录里把 Claude 初始化起来。咱们先给它一个起步 prompt,让它把这个 agent 搭出来。我会让它给我搭一个针对 Arize AI Phoenix 产品的 PM agent。我可以直接把那个完整 repo 的 URL 链接放进这里,这样它就能拿到我让它搭什么的确切上下文。然后我接着问:我想让它拿到哪些上下文?所以——把最近的 GitHub discussions 拉下来,把最近的 release 全拉下来,再看 GitHub issues。我这里会先一点一点来,先只从一个来源——也就是 GitHub——开始喂上下文。等我们把这套规模做大,你就可以像我刚说的那样,把 Gong 转录、产品分析、Slack 对话、Discord 频道都加进来,什么都能塞进去。我首先想做的,是先给这些 issue 和 discussion 按优先级打分。也就是先搞清楚:我们想让它去看、去做的这些东西到底有多重要。要看的维度比如 bug 还是 feature、大家给的 reaction、评论,还有——我确实想让它考虑时效性(recency)。这些都是我让这个产品品味 agent 去看、去权衡的东西。然后调用 Claude——我这里可以更具体,可以说调用 Claude Opus,或者我想要的任何模型。所以调用 Claude,配上……我甚至可以让它去做某种 prompt caching,这样它不用每次跑这个闭环都重新把 issue 全拉一遍。不过为了一开始简单点,我现在就只是调用 Claude,然后让它写出一份 markdown 格式的 PM 报告,里面包含——
[11:55] Aparna
you know, that has as the output the top pain points feature asks and and themes. Order this by P0 to P3 priority. So, this is basically going to be like initial starter prompt for me to actually build this product taste. I can get super, you know, typically what I like to do is uh be really thoughtful about the plan that I'm giving my agent so that it you know, it's not just going off of nothing, but you know, there's also times where you'll just have it go off, build something, and then you're iteratively giving it feedback, and that's totally also okay. So, um and then I'll just say here, use my GitHub token and my Anthropic API key. So, let's see what it can come back just with that. Super simple. Um while this is going and kind of doing its thing in the background, what I'm actually going to show you, you can see it's going to interrupt and ask a ton of questions as we go through this, but what I'm actually going to show you all is just a a simple one I built right before this and see if we can get the one we're building right now to just match up um and and see how how close we can get in just an hour here. Okay, so this is basically a PM agent that is already built out and already kind of um you know, we've had tracing set up and is sending to Arize already. And I'm just going to open one of these so I can show you all kind of what it looks like here, but this PM agent is these are the traces of our actual PM agent. And for those of you who are like, "What's a trace?" Like that's that's, you know, new concept to to understand. Um you can think about a trace really just as um it is the step-by-step playback of what this agent actually did. In this scenario, this agent is first going ahead and pulling back GitHub discussions, it's pulling back
嗯,输出里包含最主要的痛点、feature 诉求和各种主题。按 P0 到 P3 的优先级排好序。所以这基本上就是我用来真正搭出这个产品品味(agent)的初始起步 prompt。我可以做得很细——我通常喜欢的做法是,认真想清楚我给 agent 的那份计划,让它不是凭空乱跑;但也有些时候你就直接让它跑、让它先搭点东西出来,然后你再一轮一轮给它反馈,那样也完全 OK。然后我会在这儿加一句:用我的 GitHub token 和我的 Anthropic API key。咱们看看就凭这些它能给我返回什么,超级简单。趁它在后台跑、做它的事,我先给你们看——你会看到它跑的过程中会打断、问一大堆问题——我先给你们看一个我在录这期之前刚搭好的简单版本,看看能不能让我们现在正在搭的这个最后对得上,看看在这一个小时里我们能做到多接近。好,这基本上就是一个已经搭好的 PM agent,而且我们已经把 tracing 配好了,它已经在往 Arize 发数据了。我打开其中一条给你们看看大概长什么样——这个 PM agent,这些就是我们这个真实 PM agent 的 traces。对那些会问『trace 是啥?』、觉得这是个全新概念的同学——你可以把一条 trace 理解成:它就是这个 agent 实际做了什么的、一步一步的回放。在这个场景里,这个 agent 先去把 GitHub discussions 拉回来,再去拉
[14:02] Aparna
the GitHub issues, it's figuring out what are all the releases that were recently released. And then it's going through and it's actually looking at every single issue that is inside of that project and it's actually consuming all of these and coming up with the score of how important each of these issues that it's raised are. As a product person, this is kind of the first thing you need to understand is like how important are all of these asks that are coming from your users? What is the pain that it's solving? Um and so the first thing I'm just asking it to do is figure out, well, can you score basically how important is each one of these asks that are coming back from for this project. And what I'll actually do, you know, as as it scores, I won't actually have an eval that will evaluate how good was the score that my PM agent actually came up with. And is it accurate or inaccurate based off of you know, the context that I have around how I want to prioritize bugs, how I've historically prioritized feature requests. And so I actually want to write an eval that will help teams kind of evaluate the quality of this initial PM agent that we've built. Mhm. Go back and check on our agent here and see how far we've gotten. Um so still kind of thinking
GitHub issues,再搞清楚最近发布了哪些 release。然后它会一条一条地过这个项目里的每一个 issue,把它们全都消化一遍,再给它筛出来的每个 issue 算出一个『有多重要』的分数。作为一个产品人,这其实是你需要理解的第一件事:从你的用户那儿冒出来的这些诉求,到底有多重要?它解决的痛点是什么?所以我让它做的第一件事就是:能不能帮我打个分,看看这个项目里返回回来的每一条诉求各自有多重要。接下来我实际会做的是——它打分的同时,我会写一个 eval,去评估我这个 PM agent 给出的分数到底好不好、准不准,依据是我对『该怎么给 bug 排优先级、历史上怎么给 feature request 排优先级』这些上下文的判断。所以我其实想写一个 eval,来帮各个团队评估我们刚搭出来的这个初版 PM agent 的质量。嗯。回头看看我们这边的 agent,看进展到哪儿了。嗯,看起来还在想。
[15:33] Aakash
somebody's setting up this repo correctly, like basically you created a new GitHub repo. You gave it your Anthropic API key and you just And I guess to create the repo you have to log in to GitHub. Those are the main steps people have to do before this. Correct. Correct. And I'm happy to go ahead and, you know, send you guys the you know, a sample repo if you want to get started doing this yourself so that you can follow along with a project of your choice. But in this case you can see, great. Okay. So it's gone ahead. It's actually built this agent. I'm going to go ahead and um you know, it just looks like it's updating what the um what the Okay, great. So it's actually just updated. It's using my GitHub token. It's using my Anthropic API key. And now it's actually going to go ahead. It's pulled 40 discussions, 60 issues, eight releases. And now it's going to go ahead, score each item, and then based off of the score that it gives every single one of these issues, it's going to go ahead and give me a report about what the most important things to actually, you know, top pain points, feature requests, themes, what shipped, and give me a game plan that I can then use as a starting point when I come in. A really useful feature that um you know, if you you'll do this once today, but ideally you want this kind of running all the time. Kind of consistently every time someone adds a new bug report, adds a new issue, it's kind of always doing this. So what you can do is actually just say, can you run this in a loop? Can you run this in a loop? And you can specifically say using the Claude loop kind of skill. Um This is really awesome because what Claude does is that it spins up essentially a cron job. Um well, what's a cron job? It's basically you asking
前提是有人把这个 repo 配置对了,对吧——基本上你建了一个新的 GitHub repo,给了它你的 Anthropic API key,然后你就……我猜要建这个 repo 你得先登录 GitHub。这些就是大家在这之前要做的主要步骤吧。没错,没错。我也很乐意发给大家一个示例 repo,如果你想自己上手做,就可以拿你自己挑的项目跟着一起做。不过在这个例子里你可以看到——好,行,它跑完了,它真的把这个 agent 搭出来了。我点进去看看,嗯,看起来它在更新那个……好,行,更新好了。它在用我的 GitHub token,在用我的 Anthropic API key。现在它要开始干活了——它已经拉了 40 条 discussion、60 个 issue、8 个 release。接下来它会给每一项打分,然后根据它给每个 issue 打出来的分数,给我出一份报告,告诉我真正最重要的东西是什么:最主要的痛点、feature requests、主题、已经上线了什么,再给我一份可以当起点用的行动方案,等我进来就能直接接着干。还有个特别有用的功能——你今天会跑这么一次,但理想状态是你想让它一直在跑。每次有人新提一个 bug report、新加一个 issue,它就一直在做这件事。所以你能做的,就是直接说:能不能让它跑成一个 loop?能不能跑成一个循环?而且你可以特别指明用 Claude 那个 loop 的 skill。这个特别赞,因为 Claude 会做的事,本质上是它会起一个 cron job。那 cron job 又是啥?基本上就是你让
[17:34] Aparna
Claude to be able to run some type of workflow that you do every day in a loop. Um And so in my case every day, every hour, every You could set this to every 5 minutes if you wanted to. It'll go ahead and um it'll go ahead and actually uh run this loop every you know, however cadence you set so that it actually does your job every hour you have the latest report of what you should be prioritizing for your agent. So, let's go ahead. Oh, it looks like I need to go ahead and set my GitHub token. So, give me 1 second and let me do that. Um and then we can actually go ahead and run this agent and you can watch it live. So, this is actually going ahead and running my Phoenix PM agent. Um I'm going to show you guys how to do this so that uh you can also do it, but I've also kind of already set up traces. So, what does that actually mean? Tracing is the way for teams to actually get visibility into everything these agents are doing. This is kind of a really hard uh you know thing to debug because Claude is spinning off a bunch of different things and and running this in a loop and you might not always know, you know, if it comes back with slop or comes back with something great, you know, how do I go and improve it or how do I go and figure out how it did that? And so, tracing is a really awesome way to understand what your agent's doing. Today, what I'm going to actually show you is that you know you know, I'd say tracing used to be really hard. Yeah, you had to kind of go call your engineering partner to have to go and set up tracing. I think with AI, it's probably never gotten easier to do this. So, what we have is essentially skills. Uh we've released a kind of a series of call it skills that you can actually just give to your coding agent.
你让 Claude 能够把某种你每天都要做的工作流跑成一个循环。所以在我这个例子里,每天、每小时,每——你要愿意的话甚至可以设成每 5 分钟。它就会按你设的节奏去把这个 loop 跑一遍,相当于替你把活儿干了,每个小时你都有一份最新的报告,告诉你该给你的 agent 优先做什么。那咱们来跑一下。哦,看起来我得先把我的 GitHub token 设好。稍等我一下,让我把它弄好。然后我们就可以真正把这个 agent 跑起来,你能看着它实时跑。这就是真的在跑我的 Phoenix PM agent。我会教大家怎么做,这样你也能自己跑;不过我这边也已经把 traces 配好了。那这到底是什么意思?Tracing 就是让团队真正看清这些 agent 在干的每一件事的方式。这其实是个很难调试的东西,因为 Claude 会同时分出一堆不同的事来跑、还是在一个 loop 里跑,你不一定总能知道——如果它返回的是一坨垃圾(slop),或者返回的是个特别好的结果,我到底该怎么去改进它、怎么搞清楚它当时是怎么做出来的?所以 tracing 是一个特别棒的、用来理解你的 agent 在干什么的方式。今天我实际要给你们看的是——我会说 tracing 以前真的很难。对,你基本得去找你的工程搭档,让人家帮你把 tracing 配起来。但我觉得有了 AI,做这件事大概从没像现在这么简单过。所以我们做的,本质上就是 skills——我们发布了一系列、姑且叫它 skills 的东西,你可以直接把它丢给你的 coding agent。
[19:44] Aparna
This is kind of a set of Arise skills. Um you just go in, install NPX skills add. I'll show you. We'll go ahead and do this. But once you actually add this, you can just ask Claude Code to go ahead and instrument the entire agent that we asked it to go built right now. You're looking here at a whole bunch of different skills. One of them is the Arise instrumentation skill. For those of you who are curious, it's literally just in English telling what Claude Code should do to actually send trace data over to Arise. Um it makes it super easy. I'm going to show you. It's going to feel super magical and you're not going to need to wait for your you know, you're not going to need to wait for your engineering partner to have to go and do all of this lift to go get data uh from your agent to your observability platform. So, let's go do this. Um what we're going to do actually is from here I'm going to say, "Can you help me instrument this agent?" Um so, I'm going to go ahead and actually uh ask it to instrument this agent. So, what this is actually going to do is call the arise kind of instrumentation agent. Um, so you can see here, sorry, the instrumentation skill that we just talked about. So, it's going ahead, it's calling the skill. This instrumentation skill will actually first look at the code base and understand how is this agent built? What's actually calling LLM calls? Uh, what's actually calling the tool calls? And it'll go ahead and it'll figure out kind of, you know, this case, the language that it was written in is Python, the LLM provider was Anthropic. Here's the library to go use. Here's what it's actually going to go do to set up the different calls. And it says, "Cool, everything is already wired up. Sending to Arise." And is there anything else specific
这是一套 Arize 的 skills。你只要进去,运行 NPX skills add 安装一下,我给你们演示。我们这就来跑一遍。但你一旦真的把它加上,就可以直接让 Claude Code 去给我们刚才让它搭出来的那个 agent 做全套的埋点(instrument)。你们现在看到的这里有一大堆不同的 skill,其中一个就是 Arize 的埋点 skill。对那些好奇的人来说,它其实就是用英文写明 Claude Code 该怎么把 trace 数据发送到 Arize。它让这件事变得超简单,我演示给你们看。会有一种很神奇的感觉,而且你不用再等了——你不用再等你的工程师搭档去做完所有这些重活,才能把 agent 的数据接到你的 observability 平台。那我们就来做。我们要做的是,从这儿开始我会说:「你能帮我给这个 agent 做埋点吗?」我这就让它去给这个 agent 埋点。它接下来会去调用 Arize 那个埋点 agent——抱歉,是我们刚说的那个埋点 skill。所以它现在就在调用这个 skill。这个埋点 skill 会先去看代码库,搞清楚这个 agent 是怎么搭的、到底是哪里在发 LLM 调用、哪里在发 tool 调用,然后它会去判断——这个例子里,代码是用 Python 写的,LLM 提供方是 Anthropic,该用哪个库、该怎么去把各种调用配好。然后它就说:「好,一切都已经接好了,正在往 Arize 发送。」还有什么具体想改的吗?
[21:57] Aparna
you'd like to go change? So, now, let's go ahead and just see run my agent. Um, see if it sends recent traces, and I should be able to go pop over to the platform, my observability platform, and go look at traces. We'll see if there's ones that are going to show up right now from my recent run. But, it should go ahead and actually start streaming in traces from uh, the last There we go. This is everything from the last 15 minutes that's just showing up here. And so, um, you kind of basically get a way to do all of this, you know, and it figures out everything from here's the individual LLM calls, here's the actual, you know, tool calls that were made, here's the, you know, it had to go and fetch stuff from GitHub, it had to go score every single individual LLM call, and then it finally had to come back with that report that I asked for, which was, "What are my top pain points? What are my top feature requests? Uh, what was already kind of shipped, and so you can see here it's giving me an executive summary, my top pain points, and kind of the the things that it scored really, really highly for me to go and prioritize for my product. And so, literally, I didn't open any IDE, I didn't open anything, um I literally just asked Claude code to build me an agent, gave it a really good prompt, and and then I asked it, you know, kind of what I was hoping for, and then I asked it go instrument my agent with the Rise using the skill, and boom, now I have visibility into my agent. Um everything's probably not going to be perfect, and I can probably already guarantee you that that it's not going to be perfect, but what we can do is actually start using this as a way to understand, well, um how would I improve this agent? What I'm going to actually show you
你想改的?那现在我们就来跑一下我的 agent,看看它会不会发送最近的 traces,然后我应该就能切到平台、切到我的 observability 平台去看 traces。我们看看现在能不能从我刚才那次运行里冒出一些来。它应该会开始把 traces 流式地灌进来——就是上一次……来了。这就是过去 15 分钟里所有的内容,正显示在这儿。所以你基本上就拿到了一种能把这一切都做好的方式,它会把所有东西都梳理出来:这是一个个的 LLM 调用,这是实际发出的 tool 调用,它要去 GitHub 拉东西,要给每一个 LLM 调用打分,最后还得给我返回我要的那份报告——也就是:「我最大的痛点是什么?最多人要的功能是什么?哪些已经发布了?」你看这里它给了我一份执行摘要、我的主要痛点,还有它打分特别高、建议我去优先处理的那些东西。所以说,我真的没打开任何 IDE,什么都没开,我就只是让 Claude Code 帮我搭一个 agent、给了它一个很好的 prompt,然后问它我想知道的东西,再让它用那个 skill 给我的 agent 接上 Arize,砰,现在我就对我的 agent 有了可见性。当然,不可能样样都完美,我几乎可以打包票它不会完美,但我们能做的是,把这个当作一个起点去理解:那我该怎么改进这个 agent?我接下来要给你们看的是……
[24:02] Aparna
right now is actually an in-product agent that we've built called Alex. Alex is an agent that sits inside of uh our our kind of product, and you can ask all sorts of questions like, um you know, help me figure out the common types of issues that uh are coming up. And so, this will actually go through, it'll look across the data like the inputs and the outputs, and it'll start to surface up common types of issues that users are asking from my traces. Um I can use this to actually first figure out what types of evals should I actually be running on top of my agent. Um and the reason why that's interesting is that you're you're starting your evals from a place of actually looking at your traces, looking at your errors, and trying to understand well, did it actually score some things correctly? Did it not score some things the way that I would have prioritized things? You know, how many times have you had someone on your team kind of say something was super important, super priority, but you wouldn't have given it that high of a ranking for yourself. And so, the the next thing that I really want to show is really for teams is how can you use Cloud Code actually help you figure out a baseline kind of eval for these agents that you're building. You can have it start just build a baseline eval and use that to actually iteratively improve your eval so that you're not starting from complete scratch. So, what we can do here, you can do this in our product. You can also do this uh you know, kind of you can also do this using Cloud Code again. So, kind of in the theme of today, I'll actually do this using Cloud Code um and show you how you can set up evals directly from your terminal. But, you're going to see here, once I have the traces centralized, I
我现在要给你们看的,是我们做的一个内嵌在产品里的 agent,叫 Alex。Alex 是一个就坐在我们产品内部的 agent,你可以问它各种问题,比如「帮我找出现在反复出现的常见问题类型」。它会去把数据跑一遍——比如输入和输出——然后开始从我的 traces 里浮现出用户提问中常见的问题类型。我可以先用它来搞清楚:我到底该在我的 agent 上跑哪些 evals。这之所以有意思,是因为你的 evals 是从一个真正去看自己的 traces、看自己的错误的地方起步的,去理解:它到底有没有把一些东西打分打对?有没有把一些东西按我会优先的方式来排序?你想想,多少次有团队里的人跟你说某件事超级重要、超级优先,可换成你自己根本不会给它那么高的排名。所以接下来我特别想给团队展示的是:你怎么用 Claude Code 帮你给这些你在搭的 agent 搞出一个基线的 eval。你可以让它先建一个基线 eval,再拿它来不断迭代改进你的 eval,这样你就不是完全从零开始。我们这里能做的——你可以在我们产品里做,也可以同样用 Claude Code 来做。所以按今天的主题,我就用 Claude Code 来做,给你们演示怎么直接从终端里把 evals 配起来。你们待会会看到,一旦我把 traces 集中起来,我……
[26:03] Aparna
can actually ask can you suggest uh a good eval for my agent. I want it to Uh and you know, I could just start with that. Can you suggest a good eval for my agent? Um let's see what it comes back with. What this will actually do, it'll call um the skill, the evaluator skill, um that actually looks at uh looks across the traces and suggests kind of um there we go. Okay, so, looks across the skill and it suggests, okay, well, these are kind of three evals that you might want to do. There's report groundedness, checks whether the quotes, the issues in the final PM report are grounded in the actual data fed in. Um it runs kind of across everything. So, it's almost, you know, I think about this almost like an eval on the final report that was created. You could do an eval on priority alignment, checks whether the P0, P1 kind of in the report matches the top scored issues from kind of what you're expecting. Um or something around report actionability. Okay, well, I could do these, but these are all things that are kind of looking across almost like the end product. What I actually want is um is something different as a PM. Um what I want is actually to look at every single I want to get a little bit more granular in the beginning and start to understand for every single issue that this kind of for every single one of these issues here, did it actually give it a right score? Like in this case, it said that you know, it gave it a priority of a three. In this case, I don't know, let's let's pick another set of them. This case, it gave it a zero. It said, you know, this integration is not that important. Um it gave this privacy question a three. And so, there's kind of all of these it's kind of making up these priorities. And I actually wanted it to
就可以直接问:「你能给我的 agent 推荐一个合适的 eval 吗?」我想让它……其实我可以就这么开头:「你能给我的 agent 推荐一个好的 eval 吗?」我们看看它会返回什么。它接下来会去调用那个 skill,evaluator skill,它会去看一遍 traces,然后给出建议……来了。好,它看了一遍,建议说:好,这里有大概三个你可能想做的 eval。一个是报告 groundedness(事实依据),检查最终那份 PM 报告里的引述、那些问题是不是真的有喂进去的实际数据作依据。它会把所有东西都跑一遍。所以这个我几乎可以理解成:针对生成出来那份最终报告的 eval。你还可以做一个优先级一致性的 eval,检查报告里的 P0、P1 是不是和你期望的、得分最高的那些问题对得上。或者做一个围绕报告可执行性的。好,这些我都能做,但它们都属于那种「站在终点、看成品」的视角。而我真正想要的是不一样的东西。作为一个 PM,我想要的其实是去看每一个……我想在最开始就更细一点、更颗粒化一点,去理解每一个问题——对它这里列出的每一个问题——它打的分到底对不对。比如这个例子里,它给了一个优先级三。这个里,我看看,咱们再挑一组。这个它给了零,它说这个集成没那么重要。它给了这个隐私问题一个三。所以这一堆优先级它基本上是自己编出来的。而我其实想让它……
[28:13] Aparna
first just evaluate is the score that it's attaching to kind of determine how important these issues are, is that actually something that I would have said by myself. So, I actually wanted to run something like uh you know, a priority score, a priority kind of eval on um you know, is the score that it's actually saying how important these GitHub issues are, are they actually accurate based off of um how I want to weight them. So, let's go back to Cloud Code. I can actually just ask it to help me come up with an with a way to eval this. Um and this is very normal where you're kind of doing this back and forth with Claude and you're actually asking it to to go back and repeat yourself and um you know get really specific about what you want. So, in this case I can ask, "Can you help me build an eval um to evaluate uh if each issue is actually uh scored correctly?" Um or it's each issue's priority, maybe, is a good way to say this. Issue choose each issue's priority is actually scored correctly. I think that's option two, right? Priority alignment? Yeah. Yeah, this is Oh, um well, this is this is slightly more about at the cuz it looks like it's checking at the end in the very report if the top scored issues are kind of what I would have picked. Um but what I'm looking for is something slightly more nuanced, which is not just the top issues, but every single kind of individual issue is actually um given its appropriate kind of weight. So, Mhm. it's kind of giving me this like priority accuracy evaluator. Um so, it'll go ahead, it'll create a way to to run this evaluator on top of the actual traces. Uh in this case it's already picking one that I've actually already created um to do this, just to show you guys kind of how this works. Um but it'll kind of suggest, "Hey,
先就去评估一件事:它给这些问题贴上的、用来判断这些问题有多重要的那个分数,是不是我自己也会这么打。所以我其实想跑一个类似优先级打分、优先级 eval 的东西,看它给出的这些 GitHub issue 有多重要的分数,按照我想怎么给它们赋权来看,到底准不准。那我们回到 Claude Code。我可以直接让它帮我想出一个评估这件事的办法。这种来回反复其实非常正常——你就是在跟 Claude 一来一回,让它回过头来重说一遍,让你把你想要的东西说得特别具体。所以这个例子里我可以问:「你能帮我建一个 eval,来评估每个 issue 是不是真的被打对了分?」或者说,每个 issue 的优先级——也许这么说更好——每个 issue 的优先级是不是被准确地打了分。我觉得这就是第二个选项吧?优先级一致性?对吧。对。这个嘛……嗯,这个稍微更偏向于……因为看起来它是在最后那份报告里检查,得分最高的那些问题是不是我会挑的那些。但我要找的是更微妙一点的东西——不只是排名靠前的那些问题,而是每一个单独的 issue 是不是真的被赋予了它应得的权重。所以呢……嗯,它给我搞出了这么一个像是「优先级准确度」的 evaluator。它接下来会去创建一个办法,把这个 evaluator 跑在实际的 traces 上。这个例子里它已经挑了一个我之前其实已经建好的,就为了给你们演示这是怎么运作的。但它会建议说:「嘿,
[30:31] Aparna
there's this eval you've already created which is kind of doing this like row level, issue level kind of priority." Um and then it's actually going to use this to go and run it on top of those kind of traces. So, in this case it's saying, "Hey, it's running it on older data. Do you want to go ahead and run it from today's issues, like the new issues that you just grabbed from today?" So, it'll go ahead and start running it on the newer spans. Um and you can see here every single kind of GitHub issue that has come in, it's going to go ahead and give it a score of how important it actually is. Um and then it'll evaluate whether that score that it was given was actually an appropriate eval or not. I hope you're enjoying today's episode. Are you interested in becoming an AI product manager making hundreds of thousands of dollars more joining Open AI or Anthropic? Then you might want to do a course that I've taken myself, the AIPM certificate ran by Open AI product leader Mick Dad Jaffer. If you use my code and my link, you get a special discount on this course. It is a course that I highly recommend. We have done a lot of collaborations together on things like AI product strategy. So, check out our newsletter articles if you want to see the quality of the type of thinking you'll get. One of my frequent collaborators, Pavel Hearn, is the Build Labs leader. So, you're going to live build an AI product with Pavel's feedback if you take this AIPM certificate. So, be sure to check that out. Be sure to use my code and my link in order to get a special discount. Here's the dirty secret about prototyping. You spend 2 weeks building a prototype. You validate your assumptions. Engineering loves the direction. Then what happens? You throw the whole thing away. Bolt changes this
你之前已经建过这么一个 eval,它做的就是这种行级别、issue 级别的优先级判断。」然后它会真的拿这个去跑在那些 traces 上。所以这个例子里它说:「嘿,它现在跑的是比较旧的数据。你要不要直接拿今天的 issue 来跑,就是你今天刚抓下来的那些新 issue?」于是它就开始在更新的那些 span 上跑。你们可以看到,进来的每一个 GitHub issue,它都会去给它打一个「到底有多重要」的分,然后再去评估它给的那个分到底是不是一个合适的评估结果。希望你们喜欢今天这一集。你有兴趣成为一名 AI 产品经理、多挣几十万美元、加入 OpenAI 或 Anthropic 吗?那你也许会想上一门我自己也上过的课程——由 OpenAI 产品负责人 Mick Dad Jaffer 主办的 AIPM 证书课。用我的优惠码和链接,你能拿到这门课的专属折扣。这是一门我强烈推荐的课。我们一起做过很多合作,比如 AI 产品策略方面的。所以如果你想看看你将得到的那种思考的质量,可以去看我们的 newsletter 文章。我的常合作者之一 Pavel Hearn 是 Build Labs 的负责人,所以你要是上了这门 AIPM 证书课,会在 Pavel 的反馈下实打实地搭一个 AI 产品出来。一定要去看看,记得用我的优惠码和我的链接来拿专属折扣。说说原型设计里那个不能说的秘密:你花两周搭了个原型,验证了你的假设,工程团队也很喜欢这个方向。然后呢?你把整个东西扔掉。Bolt 把这个彻底改变了……
[32:11] Aparna
completely. When you prototype in Bolt, you're not building throwaway mockup. You're building real front-end code that integrates with your existing design system. So, when you hand it to engineering, they don't throw it away. They ship on top of what you've built. I use Bolt every single day. I host my Land PM job cohort on it. And honestly, I'm up till 2:00 a.m. some days just vibing in the tool, having fun, and building. That's when you know a product is good. When you're using it past midnight, not because you need to, but because you want to. Check out Bolt at bolt.new link in the show notes. I used to think [music] I had a retention problem. Turns out I had a messaging problem. I was sending the same onboarding emails to every [music] new user whether they activated on day one or never logged in again. I had no idea who was slipping or why. Customer.io changed [music] that. Every message I send is now based on what users actually do in the product. [music] Someone hits a key activation moment, they get nudged to the next one. Someone goes quiet, they get a different [music] path entirely. Their AI agent makes it fast. I describe the campaign I want and it builds the full journey [music] for me. Triggers, timing, copy, even branching logic. And when I want to know how something is performing, I just ask [music] the agent directly and it tells me what to do next. They also have an MCP server, which means AI tools like Claude can see directly what's happening in your Customer.io workspace. Your segments, your customer data, your attribution, all of it. So instead of explaining [music] your business context every time you need help, Claude already knows it. Notion used Customer.io to personalize their onboarding and hit nearly 50% open [music] rate. Improved
……彻底改变了。当你用 Bolt 做原型时,你做的不是用完就扔的 mockup,你做的是真正的前端代码,会和你现有的设计系统对接。所以当你把它交给工程团队,他们不会扔掉,而是直接在你做的东西上接着往上发布。我每天都在用 Bolt,我的 Land PM 求职训练营就托管在它上面。说实话,我有些日子会熬到凌晨两点,就在这个工具里「vibe」、玩、做东西。一个产品好不好,看这个就知道了——当你过了半夜还在用它,不是因为非用不可,而是因为你想用。去 bolt.new 看看 Bolt,链接在节目说明里。我以前以为我有个留存问题,结果发现我有的是个信息触达问题。我当时给每一个新用户发的都是同一套 onboarding 邮件,不管他是第一天就激活了,还是从此再没登录过。我完全不知道谁在流失、为什么流失。Customer.io 改变了这一点。我现在发的每一条消息,都是基于用户在产品里实际做了什么。有人触达了某个关键的激活时刻,他就会被引导去下一步;有人不吭声了,他就会走上一条完全不同的路径。它的 AI agent 让这一切都很快。我描述一下我想要的活动,它就帮我把整套旅程搭出来——触发条件、时机、文案,甚至分支逻辑都有。等我想知道某件事表现得怎么样时,我直接问那个 agent,它就告诉我下一步该做什么。他们还有一个 MCP server,意味着像 Claude 这样的 AI 工具能直接看到你 Customer.io 工作区里正在发生的一切:你的分群、你的客户数据、你的归因,全都能看到。所以你不用每次需要帮忙时都重新解释一遍你的业务背景,Claude 已经知道了。Notion 用 Customer.io 把他们的 onboarding 做了个性化,open rate 接近 50%,还提升了……
[33:38] Aparna
conversion by 6 to 7% with localized campaigns and pushed open rates up another 20% through [music] AB testing. The idea is simple. Customer.io helps you deliver more impact from every message you send. If you're a PM or founder and your onboarding is still one size fits all, try Customer.io at customer.io. I'm keen to see what evals it creates. So I guess the traditional sort of evals teaching literature is all about like you finding production traces that you feel like there was an error. So I guess
……用本地化的活动把转化率提升了 6 到 7%,还通过 A/B 测试把 open rate 又往上推了 20%。道理很简单:Customer.io 帮你从发出的每一条消息里榨出更大的影响力。如果你是 PM 或创始人,而你的 onboarding 还是一刀切的,去 customer.io 试试 Customer.io 吧。我很好奇它会创建出什么样的 evals。那我想,传统那套讲 evals 的教学文献,基本都是讲你怎么去找那些你觉得出过错的生产 traces。所以我想……
[34:13] Aakash
Right. that line of thinking would say you'd go to the trace dashboard in a rise. You'd look at those priorities, you'd say, "Oh, this is a zero, but this really should have been a four." Right.
对。那种思路会说,你应该去 Arize 里的 trace 仪表盘,看那些优先级,然后说:「哦,这个是零,但其实应该是个四。」对吧。
[34:25] Aakash
And then you'd pick up like 50 of those errors, then you'd group them and say like, "Okay, these are the 10 errors that it does." So, is that Are we trying to replicate that process, but have Claude Code basically do it itself? Is that what we're doing here? Exactly. Exactly. So, basically what Claude Code is doing is it has access to all of the traces in our eyes because the skill was basically it can go and call an API um and it can kind of share what it's doing under the hood. Um so, we can talk about it cuz it does feel a bit a little bit magical. When when we kind of just talk talk through it. Um so, give me 1 second. Let me kind of share the secret sauce of kind of what's happening here. Under the hood, all of these skills were actually calling uh APIs. And specifically the APIs that skills um tend to call is that what we've realized is that these coding agents are really good with command line or CLI interfaces. So, what it's doing is basically under the hood calling and fetching all of the traces. And you know, you've seen kind of Hamel and Shreya tell you, "Hey, go through line by line, look at where the individual traces failed." Um that is totally, you know, a great way to do this. You can of course go in and get started and start doing annotations and start doing, you know, like, "Did it actually answer the question? Is this you know, you can write free form text." And just write free form text about like, you know, what was good, what was wrong about this. Um it's absolutely a great way to do that. Um I'm also someone who I love to see if Claude Code can help me cut some of that time and surface up some insights for me. And so, what I'm actually doing here is trying to understand just with Claude Code and if I can give it access to my
然后你会拣出大概 50 个这样的错误,把它们归个组,说:「好,这就是它会犯的 10 类错误。」那么——我们是不是想复刻这个过程,只不过让 Claude Code 自己来做?我们这儿做的是不是这个意思?(Aparna:) 正是,正是。基本上 Claude Code 在做的事,就是它能拿到我们 Arize 里所有的 traces,因为那个 skill 基本上就是它能去调一个 API,还能把它在底层做的事分享出来。所以我们可以聊一聊,因为它确实感觉有点神奇——当我们就这么把它讲一遍的时候。那给我一秒钟,我来分享一下这背后到底发生了什么的「秘方」。在底层,所有这些 skill 其实都在调 API。具体来说,这些 skill 倾向于调用的那种 API——我们发现的是,这些编码 agent 特别擅长跟命令行、跟 CLI 接口打交道。所以它在底层做的,基本上就是去调用并把所有 traces 拉下来。你们也听过 Hamel 和 Shreya 跟你们说:「嘿,一行一行地过,看哪些单独的 trace 失败了。」那当然是一种很好的做法。你当然可以进去、上手开始做标注,开始去做——「它到底有没有回答这个问题?」你可以写自由文本,就写自由文本,写写这条 trace 哪里好、哪里不对。这绝对是一种很好的做法。但我也是那种喜欢看看 Claude Code 能不能帮我省点这种时间、帮我浮现一些洞见的人。所以我这里其实是想搞清楚:光靠 Claude Code、如果我能给它接入我的……
[36:27] Aparna
spans and my traces, like what are some insights from this that I should have to go and learn you know, help me go and tell me what's wrong with my agent. And sometimes it's you know, just being super honest, like sometimes it might not come back with something amazing as your first eval, but what I typically like about it is that it gives me a place to actually start thinking about problems and start thinking about areas of of improvement. So, in this case, I've gone ahead uh and created this like priority accuracy. Like priority accuracy eval, and it's running it's now running it's run across all of my new spans, and I can go in here and just say show me everything where the label is actually inaccurate. Where Claude code thinks that the priority, you know, you can see the scores here, the priority that it's come up with is actually wrong, and why is it wrong? And you know, this is probably something that you're going to hear all the time from folks who do evals is was my eval wrong or was my agent wrong? And you will definitely have scenarios, and there's a whole process that Hamel and Trey actually talk a lot about, which is aligning your evals so that your evals uh are grounded in that kind of human feedback. What I'm sharing is kind of a way right now of can you start it's almost like you know, can you start with the vibe eval and then modify it and improve it so that it becomes something that you you can trust and and go from. And uh you can do either approach. You can go through the you know, actual coding approach, surface up all the issues, have the human in the loop, uh and you know, identify categories of pain. Um but as a product person, you might already know what types of things you definitely want to catch. For me, what I want to catch is
……我的 span 和 traces,那从这里面能挖出哪些我本该去学的洞见——帮我去看看、告诉我我这个 agent 哪里有问题。有时候吧,说句大实话,它作为你的第一个 eval 可能并不会返回什么惊艳的东西,但我通常喜欢它的地方在于,它给了我一个真正可以开始思考问题、开始思考改进方向的起点。所以这个例子里,我已经去建了这么一个「优先级准确度」的 eval,它正在跑——现在已经在我所有新的 span 上跑完了——我可以进到这里来,直接说:把所有标签其实不准确的地方都给我看一下。就是 Claude Code 认为那个优先级——你能在这儿看到分数——它给出的那个优先级其实是错的地方,以及它为什么错。你知道,做 evals 的人多半会一直听到这么一个问题:是我的 eval 错了,还是我的 agent 错了?你肯定会碰到各种情况,而且有一整套流程,Hamel 和 Trey(Shreya)实际上聊了很多,就是怎么去对齐你的 evals,让你的 evals 扎根在那种人类反馈之上。我现在分享的,算是一种思路:你能不能先从一个「凭感觉的 eval」(vibe eval)起步,再去修改它、改进它,让它变成一个你能信得过、能拿来用的东西。这两种路子你都可以走。你可以走那条真正写代码的路,把所有问题都浮现出来,让人类介入(human in the loop),去识别出各类痛点。但作为一个做产品的人,你可能本来就已经知道你绝对想抓住哪些类型的东西。对我来说,我想抓住的是……
[38:35] Aparna
is every single issue that this agent is prioritizing, is it right or is it wrong? Is it accurate? Is it giving it an accurate score or is it not giving it an accurate score? And I can start off by saying, well, let me see if I can just have it go and create an eval to suggest kind of what that priority accuracy looks like. Um you can do it through a skill. You can also, if you do have human annotations that are built through here, it will the skill will look at those human annotations and use it to actually build you an eval as well. So, um I in this scenario didn't have any, but if I had one, it would go through and do the whole process that Hamel and Treya kind of walked through of like aligning the evals. So, it's gone ahead, it's run kind of the priority accuracy eval. It's comparing the accuracy of you know, it's something that's looking at the score that was assigned to each of the issues and it's surfacing up kind of, you know, is this an accurate score or is this not an accurate score? Again, this is just based off of, you know, a simple first pass of this eval. I am going to refine this eval now. Because this eval is completely, you know, based off of just Claude looking at my traces and trying to identify problems. Um and the whole point of this is like, how do we get this loop kicked off? This loop is meant to kind of give you a starting spot. It is not meant to be your end-all-be-all kind of you know, state for your evals or your agent. Your evals will adopt, uh will will kind of get better and your agent will get better. And that's kind of what we're showing in this workflow today is kind of how do you get started, how do you get unblocked, and then how do you do that improvement loop so we can make this better. So, in this case I kind of have like a
就是这个 agent 排出来的每一个 issue,到底对不对?准不准?它给的优先级打分准确吗?我可以先这么入手:让我看看能不能直接让它去生成一个 eval,来评估这个优先级的准确度大概是什么样。这个你可以通过一个 skill 来做。另外,如果你这里已经积累了人工标注,这个 skill 也会去读那些人工标注,并用它来帮你真正搭建出一个 eval。所以在这个场景里我手头没有标注,但如果有的话,它就会走一遍 Hamel 和 Treya 之前讲的那整套流程,也就是把 eval 对齐。好,它已经跑完了,跑的是这个优先级准确度的 eval。它在比对准确度,也就是去看分配给每个 issue 的打分,然后把结果浮现出来,告诉你这个分到底准不准。再强调一下,这只是基于这个 eval 一次很简单的初跑。我现在要去打磨这个 eval。因为这个 eval 完全是 Claude 看了我的 traces、试着识别问题之后生成的。而这一切的重点在于:我们怎么把这个循环启动起来?这个循环只是给你一个起点,它不该是你 eval 或 agent 的最终形态、终极答案。你的 eval 会逐步适应、会越来越好,你的 agent 也会越来越好。这正是我们今天在这个工作流里要展示的东西——你怎么起步、怎么破局,然后怎么跑这个改进循环,把它越做越好。所以在这个场景里,我大概有这么一个
[40:25] Aparna
very simple small eval here which is, okay, looking at the accuracy of the score, these are ones that Claude thinks are not accurate. I can actually just directly ask here like, when my priority accuracy is inaccurate, you know, what are common issues or reasons for that? So, this will actually kick off and now look at, well, what types of uh what types of uh you know, things is my PM agent not uh prioritizing correctly. So, I have my agent kind of kicking off, looking at the data, and what we're trying to do here is really go from you built an agent, you have traces set up automatically through Claude Claude, you have Claude kind of suggesting what an eval could actually look like, and now these are already scenarios that Claude thinks are are not right accurately scored. This is a great starting ground for me to say, okay, well, what can I go to understand how to go improve this agent? And you barely had to write you didn't have to write anything. You kind of had to, you know, ask Claude a couple couple things. Um so, let's go ahead and this is Alex kind of giving me suggestions of of kind of what to go do here. So, in this case there's uh whole categories of issues that it's looking at. So, there's somewhere there is a feature request scoring, there's a legacy scoring system, there's bugs priority scoring, there's low priority scoring, there's data fetch. Okay. So, there's a lot of different categories of where it's actually suggesting that my scoring might be off and it's giving me a whole bunch of spans to go look at to go debug and understand kind of what what are some actual problems that this PM or a taste agent might have in prioritizing issues that are coming. Um And what's a span exactly? It's a group of traces? A span is really an
非常简单、非常小的 eval,就是看打分的准确度,这些是 Claude 认为不准确的。我其实可以直接在这里问它,比如:当我的优先级准确度不准时,常见的问题或原因是什么?这就会启动起来,去看我的 PM agent 在哪些类型的事情上没有正确排优先级。所以我让我的 agent 跑起来、去看数据,我们在这里真正想做的是从「你搭好了一个 agent」起步——你通过 Claude 自动设置好了 traces,让 Claude 帮你提建议、给出一个 eval 大概该长什么样,而现在这些已经是 Claude 认为打分不准的场景了。对我来说这是个很好的出发点,我可以说:好,那我接下来该怎么去搞懂如何改进这个 agent?而你几乎不用写——你根本没写任何东西,你只是问了 Claude 几个问题而已。好,我们继续,这是 Alex 在给我建议接下来该做什么。在这个场景里,它在看一整批的问题类别。比如有 feature request 的打分、有 legacy 打分系统、有 bug 优先级打分、有低优先级打分、有数据抓取。好。所以有很多不同的类别,它在提示我的打分可能在哪里出了偏差,并且给了我一大堆 spans 让我去看、去 debug,去搞懂这个 PM agent(或者说品味 agent)在给进来的 issue 排优先级时可能存在哪些实际问题。那 span 到底是什么?是一组 traces 吗?span 其实是
[42:31] Aparna
individual step in a trace. So in this case what you're looking at here is this is this entire interaction where it did this whole report is what you'd call a trace. A span is a single individual step or a single individual issue that I had to go look at. Got it. Yeah. And it's weird. Isn't it a little bit weird that Claude rated like everything it did inaccurate? I uh some of them are accurate and some of them are
一个 trace 里的单个步骤。所以在这个场景里,你现在看到的这整段交互——它生成这整份报告的全过程——这就是所谓的 trace。而 span 是其中一个单独的步骤,或者说我要去看的某一个单独的 issue。明白了。是啊。不过有点怪——Claude 把它自己做的几乎所有东西都评成不准确,这是不是有点怪?嗯,有些是准确的,有些是
[43:00] Aparna
Okay, some of them are accurate. Okay. If it did, then that would probably be a good spot for you to understand, okay, well, maybe I shouldn't trust that eval from from Claude. Um and so I feel like a good eval is like you're getting some healthy percentage right but also healthy wrong so that you can make progress, right?
好,有些是准确的。好。如果它真的把全部都评成不准,那对你来说可能就是个信号了:好吧,也许我不该太信这个来自 Claude 的 eval。所以我觉得一个好的 eval 是这样的——你既能挑出健康比例的「对」,也能挑出健康比例的「错」,这样你才能取得进展,对吧?
[43:17] Aakash
100% and so you want that feedback of like I get excited when I see that evals are wrong because then it gives me a chance to know that there's improvement that could be made but when everything's wrong, then you know, it's obviously that's definitely a scenario where you need to start looking at your eval to understand what what to go improve. And when when can we do the vibe evals? When do we have to do the axio coding or can you always start from vibe evals and then a layer in axio coding talking to the agent later? So my take is that vibe evals are going to fall short very very quickly. Um it's and the reason for it is that it just doesn't have any it's not grounded on any actual human that is involved in curating that taste again of your agent. And so what you really want is something that helps you, you know, I think that it would be hard to say, "Hey, you have to go and immediately start by you know having a bunch of vibe evals and using that to evaluate your agent." Like that it just the signal to noise ratio there is going to be really really low. And so having something where you have maybe a simple thing that gets kicked off, but then now what I'm going to actually go do here is that process where I have a simple eval and I'm now going to make sure, "Okay, well, is this eval that I've created actually something that I can trust?" Um and it's not going to be. It was a one-shot eval that's out of the box. I'm going to actually go through and figure out, "Well, where do I disagree with it? Where do I not disagree with it? How do I actually" And you would do this process even if you did actual coding. Even if you did actual coding and you did individually, you know, human annotated every single span and every single issue and you were able to put together this
百分之百同意。所以你想要的就是这种反馈:当我看到 eval 判出错误时我反而很兴奋,因为这给了我一个机会,知道有可以改进的地方;但当所有东西都被判成错的,那很明显,那种场景下你就该开始回头审视你的 eval、搞清楚到底该改什么了。那我们什么时候可以做 vibe eval(凭感觉评估)?什么时候必须做那种正经的编码标注,还是说你永远可以先从 vibe eval 起步,之后再叠加正经标注、再跟 agent 对话?
[45:19] Aparna
amazing ground truth data set. Um your eval will get misaligned over time as you see more and more data. And so it is super important that uh you regularly align those evals to the data that you're actually seeing on the ground um with with your users. So um what I'm going to do right now is actually walk through a process where I've created a very simple eval out of the box. Claude just one-shotted it for me. And now I'm going to start asking, "Okay, well, is this an issue with an eval? Is this an issue with my agent?" Um and you have examples you know, in this scenario, it looks like uh bugs are bug category items using the new scoring system with category four are also commonly inaccurate. And so, it feels like there's scenarios where bugs maybe are not getting categorized or given the accurate score that I wanted to. In my world, I want bugs to always be super high because if it's a bug and a customer hits a bug, that's just a really bad experience with the product. So, I would prioritize bugs over uh you know, even new feature work. Um and so, this gives me a way to say, "Okay, well, let me go look at some examples of where the bugs are, you know, being prioritized really low." And just gives me a category of problems to start looking at and start debugging and understanding kind of how good this this agent is. Um and what you can do is, you know, for some teams, these evals end up, as you as they get really good and as they get really better, uh you can immediately ask Claude to, you know, going back to kind of using Claude code with evals, say, "Hey, go grab everything where this eval failed and suggest an improvement and go and improve improve that eval for me." I think it's uh unfair to say people aren't creating aren't create using
我的看法是 vibe eval 会很快、很快就撑不住了。原因在于它没有任何根基——它没有立足于任何真正参与的、为你的 agent 提炼品味的人。所以你真正想要的是能帮你做到这点的东西。我觉得很难说「嘿,你必须一上来就立刻搞一堆 vibe eval,拿它来评估你的 agent」,因为那样信噪比会非常非常低。所以更好的做法是:你也许先有个很简单的东西被启动起来,但我接下来真正要做的,是那个流程——我手上有一个简单的 eval,现在我要去确认:好,我创建的这个 eval 真的是我能信赖的东西吗?而它一开始不会是。它是开箱即用、一次生成的 eval。我要去走一遍,搞清楚:我在哪些地方不认同它?哪些地方我没有异议?我到底该怎么做?而且就算你做的是正经标注、就算你逐条人工标注了每一个 span、每一个 issue,拼出了一个
[47:19] Aparna
Claude to create evals. And I think that's maybe one of the pain points that I see with always saying start with actual coding is that in reality, um you will always do it, but I think it's okay to start with Claude suggesting what a good suggestion of an eval could be. And these models have gotten so good. Like, having it go through and look at your answers and suggest, "Hey, that probably is something you should flag and look at." I would trust it. I would trust it as a first pass, like go tell me what my evals should be. Yeah, that's my favorite workflow. Always start Claude generating it, but then you just give it like ruthless criticism. And I just turn on dictation mode and I'm like, "Well, you misjudged this for this reason. You misjudged this for this reason." And that's where the taste alpha that you bring can actually come back in. Totally. And I think what for me is like, how do I quickly get into that loop? Is get data in, get an eval set up, give it criticism, and let it go run on a loop. Um So, I showed earlier there's kind of the Claude loop. Um uh kind of skill that Claude has. And so, what you can actually do here is now that you have this eval, you can create a whole 'nother skill that's just like everyday go through, fetch everything that was inaccurate, and go uh that was inaccurately prioritized, and go fix and improve my agent. And you can go and create a skill that actually will then go suggest improvements to your agent from the evals that you just ran on top of this. Mm. So, you actually loop the improvement, too, not just the agent. Cuz then you get to a world of self-improvement. And that's where, to be honest, I think we're all headed is that the data that we all collect, the evals and observability is the
无比完美的 ground truth 数据集,你也照样要走这个流程。随着你看到越来越多的数据,你的 eval 会随时间逐渐失准。所以非常重要的一点是:你要定期把这些 eval 跟你在一线、在用户那边真正看到的数据对齐。所以我现在要做的,就是走一遍这个流程——我创建了一个开箱即用、非常简单的 eval,Claude 一次就帮我生成好了。现在我开始追问:好,这是 eval 的问题,还是我 agent 的问题?你看这里有例子,在这个场景里,看起来用新打分系统、被归到第四类的 bug 类目项也经常不准确。所以感觉存在这样的场景:bug 可能没被正确归类,或者没拿到我想给它的那个准确分数。在我的世界里,我希望 bug 永远是超高优先级的,因为如果是个 bug、客户撞上了它,那对产品就是非常糟糕的体验。所以我会把 bug 排在新功能之前。这就给了我一个切入点,我可以说:好,让我去看几个 bug 被排得很低的例子。它直接给了我一类问题,让我可以开始看、开始 debug,搞懂这个 agent 到底有多好。你还能做的是——对某些团队来说,随着这些 eval 越来越好、越来越精,你可以直接让 Claude(回到用 Claude Code 配合 eval 的思路)说:嘿,把这个 eval 失败的所有东西都抓出来,提个改进建议,帮我去把那个 eval 改好。我觉得,要说大家没在用 Claude 来
[49:11] Aparna
foundation for self-improving agents. And so, you get your observability in, you build an initial eval. It's a first pass. You're going to make it better. You're going to have to give it ruthless criticism to kind of make the agent better or make the eval better. And I think what teams are doing right now is they're kind of doing that iteratively. You can just create a loop that essentially starts to look at the evals, identify, you know, I I I just asked right there, "Give me the common reasons why the priority accuracy is inaccurate." Oh, it's because of the way I prioritize bugs is is uh doesn't look right. And so, what I can do, go back to my PM agent, and just say hey, go fix this issue. And then go fix the issue, ship a new agent, now go collect traces from the next rev of that agent. And so, that improvement loop can actually run inside a cloud code as a loop skill. So, that is all fine and dandy for your internal agents that are assisting you in your work. How does this all change for the AI agents in your product? You just showed us Alex. So, maybe you can go under the covers of how that worked when it's an actually a product. You're not going to be shipping self-improvement to Alex every day because you don't know it could just go off in some weird direction. So, how does Where did the human in the review loop parts come in there? Totally. I mean, there is there's still code review. There's still a human that you know, looks at every PR that is actually being put up by the self-improvement loop, but you know, maybe what I can ask you back is, but isn't that the vision? Isn't that the future that we all want to go to is that I should be able to see someone file a bug and on, you know, Alex didn't give me somebody gave a response that Alex gave a thumbs down.
创建 eval,那是不公平的。而我觉得「永远从正经编码标注起步」这个说法的一个痛点恰恰在于:现实中你最终总会用到 Claude,但我认为一开始就让 Claude 来建议一个好 eval 大概该长什么样,这完全 OK。而且这些模型已经强到不行了。让它去看一遍你的答案、然后建议「嘿,这条你大概该标记出来看一看」——我会信它的。作为初跑,我会信它,就让它先告诉我我的 eval 应该是什么样。是的,这是我最喜欢的工作流。永远先让 Claude 生成,然后你对它狂提批评、毫不留情。我就直接打开听写模式,开始念:「这条你判错了,原因是这个。那条你判错了,原因是那个。」而这正是你带来的那点品味「alpha」能重新发挥作用的地方。完全同意。对我来说关键就是:我怎么快速进入那个循环?就是——把数据弄进来、把 eval 搭起来、给它批评,然后让它在循环里跑起来。我前面展示过 Claude 那个 loop(循环)的 skill。所以你在这里实际能做的是:现在你有了这个 eval,你可以再创建一个全新的 skill,它就是「每天都去跑一遍,把所有不准确的、被错误排优先级的东西抓出来,去修」,去改进我的 agent。你可以创建一个 skill,让它基于你刚跑的这些 eval,去给你的 agent 提出改进建议。嗯。所以你把「改进」本身也变成了循环,而不只是 agent。因为这样你就进入了一个自我改进的世界。说实话,我觉得这正是我们所有人都在奔向的方向:我们收集的全部数据、eval 和可观测性,就是
[51:09] Aparna
Alex is able to immediately, and this is kind of what we're doing internally already is you're going to hear a lot more about us talking about it in the next couple of weeks. Um, but Alex is already taking that feedback, spinning up a whole debug kind of workflow, and using the eval, using the trace to debug what went wrong, and then in some scenarios, like we talked about, it's the eval that's wrong. And in that case, you know, it's a refinement on the eval. Um, and in some cases, but that's great, right? That's basically, you know, a little bit of what you hear all about the actual coding of figure out what are the reasons why that eval wasn't good and then use it to go improve that eval. And in some cases, the eval was right and it really was the agent that needed to be, you know, handle a specific scenario better. And so in that case, what we can do is just very simply um go in and do an improvement, say hey, go fix this, uh and actually go in and and improve the the agent. And so what I can do
自我改进型 agent 的地基。所以你先把可观测性接进来,再搭一个初始的 eval。它是初跑版,你要把它做得更好。你得给它毫不留情的批评,来把 agent 变好、或者把 eval 变好。我觉得现在各团队正在做的,就是这种迭代式的打磨。你完全可以建一个循环,让它本质上不断去看这些 eval、去识别问题——我刚才就在那里问了一句:「给我列出优先级准确度不准的常见原因。」哦,原来是因为我给 bug 排优先级的方式不太对。那我能做的就是回到我的 PM agent,直接说:嘿,去把这个问题修了。然后它去修、发布一个新版 agent,再去从这个 agent 的下一个版本里收集 traces。所以这个改进循环其实可以作为一个 loop skill 跑在 Claude Code 里面。那么,对于辅助你工作的内部 agent 来说,这套都挺美好的。但换到你产品里那些面向用户的 AI agent,这一切会怎么变?你刚给我们演示了 Alex。所以也许你可以掀开盖子讲讲,当它是个真正的产品时是怎么运作的。你不会每天都把自我改进直接推给 Alex,因为你不知道,它可能就跑偏到某个奇怪的方向去了。所以这里面「人在审核循环里」的部分是怎么介入的?完全理解。我是说,这里仍然有 code review,仍然有一个人去看自我改进循环提交的每一个 PR。不过也许我可以反问你一句:但这难道不正是那个愿景吗?这难道不正是我们所有人都想去到的未来吗——我应该能看到有人提了一个 bug,比如有人对 Alex 给出的某个回复点了踩。
[52:16] Aakash
That is what we want, right?
这正是我们想要的,对吧?
[52:18]
[laughter]
(笑)
[52:18] Aparna
Ideally, like it's happening like in real time across millions of users automatically. So I guess how do you do that safely? So code review is one step. What else do you need to like Where do you need to put the human in the loop? So I think there's a um there's a couple maybe places where that needs to happen. One is as the eval changes, that's also a really important step to actually having the um human kind of curate that taste of what is good and what is not good. Um so the human's kind of typically involved in eval changes, they're involved in the agent changes, um the there's a lot that's happening right now around making sure that the the skill that's actually being used to do the improvement workflows, that is one that is uh typically designed by a human. So what does that improvement skill need to look like? What is What is all of the context that it needs to have access to in order to be able to know what the improvement is. In this scenario, it might not have all the context because all I gave it was just GitHub issues. But if I could then layer in my product analytic metrics, I could layer in my um the my my traces, my actual entire traces. Um it could actually end up using that information to build its own context of what went wrong, how do I need to go fix it, um and you leverage that information is basically context for the improvement loop.
理想状态下,它就该这样发生——在数百万用户身上实时、自动地进行。所以我想问,你要怎么安全地做到这一点?code review 是其中一步。还需要什么?哪些地方你需要把人放进循环里?我觉得有几个也许必须放人的地方。第一,当 eval 发生变化时,这也是非常重要的一步,要让人去真正提炼「什么是好、什么是不好」的那种品味。所以人通常会介入 eval 的变更,也会介入 agent 的变更。另外,现在围绕「确保用来做改进工作流的那个 skill 本身是由人来设计的」这件事,也有很多工作在进行。那这个改进 skill 该长什么样?它需要拿到哪些上下文,才能知道改进到底是什么?在这个场景里,它可能没拿到全部上下文,因为我给它的只是 GitHub issue。但如果我能再叠加上我的产品分析 metrics、叠加上我的 traces、我完整的全部 traces,它其实就能用这些信息为自己构建出「哪里出了错、我该怎么去修」的上下文,你就把这些信息当作改进循环的上下文来利用。
[53:53] Aakash
Got it. So, there's human in the loop at any agent change, any eval change, but outside of that you can actually use loop commands within cloud code or whatever if you're in more production database, a real cron job, and every day or whatever cadence. And so, what are you get to work with like all of the best companies, Uber, DoorDash, you name it. What are the What is the state-of-the-art looking like for this self-improvement? How fast are people moving and how fast do they need to be moving to be competitive? I mean, I think it's going to come very, very quickly. If I'm honest with you, I think the best teams are already doing this in uh in their you know, you can call it like a radius that they're comfortable with today, but that radius is going to get bigger and bigger. Um is there you know, maybe the initial improvement is around improvements to the agent that are kind of more simpler, more around the prompts, the tools. Um does that radius then become about giving entire workflows that the agent didn't have access to do? Does it So, the radius of those changes I think is going to become uh is going to become increasingly bigger, which we're excited about. Um but it just that self-improvement loop is not going to happen without having really good um data. Really good data and really good evals. Um if you think about it and just to try to maybe take an analogy for something that's so different, but if you think about like some of the best sports players, what do they do? Like I'm talking about like the Nadals, the Federers if you're a tennis fan. Like you're Novaks. Like what they're doing is actually looking at their plays. They're looking at their previous games. They're looking and studying their behavior of what they did and using that
明白了。所以在任何 agent 变更、任何 eval 变更处都有「人在循环里」,但除此之外,你其实就可以在 Claude Code 里用 loop 命令,或者别的什么——如果你更偏生产环境的数据库,那就用真正的 cron job,每天或者按任意你想要的节奏跑。那么——你能跟所有最顶尖的公司合作,Uber、DoorDash 等等。在这种自我改进上,最前沿的状态是什么样的?大家推进得有多快,又需要多快才能保持竞争力?
[55:46] Aparna
as a way to understand what went well and what didn't go well in their games to go make improvements. This is kind of studying your plays is kind of what agents uh you know self-improving agents or self-improving harnesses have to do is they kind of have to study their own plays um to understand what did the human say was a good response or what did the human not say was a good response um and use that to actually figure out how to improve their own gameplay in some way. Um and that's what we're actually That's why the evals and the observability are kind of the the foundational layer in order for teams to actually build that self-improving loop. So, I personally have encountered PMs that I feel like are in one of three buckets and I think you have customers in all three of those buckets. So, there's the AI natives like customers you have like Handshake and the AI companies. Then there's like the digital-first companies, customers you have like Uber and Reddit and Roblox. And then there's like the normal companies who have tech arms, Pepsi, Condé Nast, normal type of companies. So, you get to you work with all three of those groups. And so, what I want to understand is usually the AI native groups they're going to be doing the quote-unquote best way or like the right way of how to do things. So, what are the AI native groups doing and specifically not just with like how they're building their evals, but the role of the PM. What is the role of a PM in a AI native company versus a company who hasn't gotten there yet and how does that company bring their PMs there? Yeah, I um I think the role of a PM is like completely changed in the last year. The role of the PM is almost like the you're the tastemaker for this product. And in order to become a really good
我觉得这会来得非常非常快。说实话,我认为最好的团队已经在做这件事了,在一个他们今天觉得舒服的「半径」范围内做,但那个半径会越来越大。也许最初的改进围绕的是对 agent 比较简单的改动,更多是关于 prompt、关于工具。那这个半径之后会不会扩展到「赋予 agent 它原本没权限去做的整套工作流」?所以我觉得这些改动的半径会变得越来越大,对此我们很兴奋。但有一点:那个自我改进循环离不开非常好的数据,非常好的数据加上非常好的 eval。你想想看——打个也许差得很远的比方——你想想那些最顶尖的体育选手,他们在做什么?比如 Nadal、Federer,如果你是网球迷的话,还有 Novak。他们在做的其实就是回看自己的打法,看自己以前的比赛,研究自己当时的行为,然后用这些
[57:41] Aparna
tastemaker, you really have to understand the outcomes of the agents. That's the especially the AI PMs where the product is the agent. The product is the agent that's being built. Um you have to spend a lot more time. You know, the AI native PMs, they are almost indistinguishable from engineers in in in some ways because they're comfortable living in Claude code. Like this entire workflow that I just showed where they're able to build even just a simple internal agent to help them do their daily tasks where they can you know, you're not doing things if you're you know, we we kind of say this internally and I think it's true is like if you're doing things the same way you were doing things last year, then you know, you haven't um you haven't caught up yet. And I think that I deeply do think that you know, if you're kind of looking at your old board of like here's my priorities and you're kind of manually scanning them and manually kind of understanding every single um you know, kind of kind of doing what you used to do, it's just different because now with the advent of kind of Claude code, I can actually have it you're not limited by how many individual meetings and Gong calls that you can personally kind of hear. You have Claude code go through and have it has access to all of these customer calls that you might have you know, never been able to consume all of by yourself, but can it help surface up the one or two that are like super critical you need to put your eyes on that because that's going to help you unlock your next 10 15 customers. And so I think in in these AI native companies what we're seeing is that the PMs were able to leverage Cloud Code to do everything from understand user data and user feedback better, surfacing that back into what
作为品味的定义者,你真的得理解 agent 的产出结果。尤其是那些产品本身就是 agent 的 AI PM——你们做的产品就是正在搭建的这个 agent,那你就得花多得多的时间在上面。你知道吗,那些 AI 原生的 PM,在某些方面已经几乎跟工程师没什么区别了,因为他们很习惯泡在 Claude code 里。就像我刚演示的整套工作流,他们能搭出哪怕只是一个简单的内部 agent 来帮自己处理日常任务。我们内部常说一句话,我觉得是真的:如果你今年做事的方式还跟去年一模一样,那说明你还没跟上。我是真心这么认为的——如果你还盯着自己那块老看板,上面写着我的优先级一二三,然后人工一条条扫、一条条去理解,基本上还在做你以前做的那些事,现在就是不一样了。因为有了 Claude code,我真的可以让它去做。你不再受限于自己一个人能亲自听几场会、几通 Gong 通话。你可以让 Claude code 去把所有这些客户通话——那些你自己一辈子也听不完的——全部过一遍,再让它帮你浮现出那一两通超级关键、必须亲自盯的,因为正是那一两通能帮你拿下下一个十个、十五个客户。所以我们在这些 AI 原生公司里看到的是,PM 能借助 Claude code 做到方方面面:从更好地理解用户数据和用户反馈,再把这些反馈反向输入到——
[59:39] Aparna
does a really good product experience look like? Uh get really close from idea to solution. So, it's not like, "Hey, I'm handing it over to an engineer." It's like they're able to effectively almost put together a plan for what that build needs to look like. Those are Those are the PMs that I think are really um going to be 10x or whatever multiplier PMs in any in any team. So, we're talking about working with those AI-native companies. You are yourself one of those AI-native companies, and you referred to this that you yourself are hiring more AI PMs than ever. So, what does the new profile look like? If I want to land an AI PM role at an AI-native company that has raised $131 million. What are the skills I should be developing? What is the depth of technical knowledge and topics I need to cover? One, I and I I've always believed this is like just the curiosity is the number That's like for me the number one most important signal. Like, this person is uh trying all the new tools. They're kind of exploring the boundaries of what they can and can't do um because that's something that you know there's there's kind of the old way of doing things is that there used to be trainings, and you'd go to these trainings, and someone would walk you through how to use a tool. Um but what if the tool is Cloud Code, and it's had you know, shipped you know, 90 features in like 30 days? Like, there is no old way of doing things where you can have like a daily training for a product that's moving that fast. And so it's kind of the onus of keeping up has become on the individual now. To actually keep up with the tools, keep up with what's changing. And if something not everything is going to be useful to what you do, but if something can give you an ability to hey, that used to take
——一个真正出色的产品体验该长什么样?从想法到方案能贴得非常近。所以不再是那种「嘿,我把它交接给工程师」,而是他们几乎能自己把这个东西要怎么搭的方案都拼出来。我觉得正是这样的 PM,在任何团队里都会成为 10 倍、或者说带各种倍数的 PM。我们刚聊到跟这些 AI 原生公司合作,而你们自己就是其中一家 AI 原生公司,你也提到你们自己现在招的 AI PM 比以往任何时候都多。那新的画像是什么样?如果我想在一家融了 1.31 亿美元的 AI 原生公司拿下一个 AI PM 的岗位,我该培养哪些技能?我需要掌握的技术知识有多深、要覆盖哪些主题?第一,我一直坚信好奇心是——对我来说这是头号最重要的信号。也就是这个人在尝试所有新工具,在不断探索自己能做什么、不能做什么的边界。因为你知道,过去那套老办法是会有培训,你去参加培训,有人手把手教你怎么用一个工具。可如果这个工具是 Claude code,而它在大概 30 天里发了 90 个功能呢?根本不存在什么老办法能让你给一个迭代这么快的产品做每日培训。所以跟上节奏的责任现在落到了个人头上,得自己去跟上工具、跟上变化。当然不是所有东西都对你的工作有用,但如果有个东西能让你——嘿,这件事以前——
[1:01:42] Aparna
me an hour and now it can take me 10 minutes. Like that is an advantage and being able to identify those and use them to your advantage is uh is deeply deeply I think it's built off of curiosity at this stage. Um I think too the other big one is it's still really important to you know care and understand like the the user and customer empathy is something that I don't think a a like the best PMs and the best product taste makers you know understand you could ask them how is that customer using the product? What's their biggest pain points? What do they you know and they would be able to rattle them off to you. And I think what's now changed is that you can actually get even deeper. You can you know customer asks for something it could have taken a week to go build that, two weeks to build that in the past. That could be delivered that day. Um if you're able to ship at that velocity. Um and so being able to get even closer and deliver to customers even faster is no longer just like a it's no longer a pipe dream. It's actually how the best products at AI natives are are are shipping right now. So 99% of people aren't in an AI native company. So they don't believe us. So I need to just confirm this is true. What you're saying is that sometimes an issue will come in. Your PMs will identify important enough. Either they will prototype or an engineer will prototype and make ready for production a feature and you guys will ship it in the same day. Yes. That is actually what's happening, guys. So, she said it herself. So, what is the role then of the PM? Like, is the PM Do PMs need to become engineers at this point? I I think that um at the AI native teams, I am seeing that that the gap between um a PM and an engineer is is indistinguishable. Um because when code is
——要花我一个小时,现在十分钟就能搞定。那就是优势,能识别出这些工具并为我所用,我觉得这在现阶段深深地、深深地是建立在好奇心之上的。我觉得另一个大点是,依然非常重要的一点是真正在乎、真正理解用户——对用户和客户的同理心,我觉得最好的 PM、最好的产品品味定义者都懂这个。你可以问他们:这个客户是怎么用产品的?他们最大的痛点是什么?他们……他们能一口气全给你报出来。而现在改变的是,你其实可以挖得更深。客户提了个需求,这事在过去可能要花一周去搭、两周去搭,现在当天就能交付——如果你能用那种速度发版的话。所以能跟客户贴得更近、交付得更快,已经不再只是个白日梦了,这就是现在 AI 原生公司里最好的产品真实的发版方式。
[1:03:57] Aparna
become so much easier to actually produce, then actually you know, this goes back to where we started today's podcast with, which is the alpha is the alpha today is product taste. So, the people that understand product taste, understand what customers want, understand how to deliver a really amazing experience, are just going to have an insane um insane velocity. Um So, PMs who can kind of go from here's the pain point, here's what I would I think is a really amazing experience and they're a triple threat where they're like, I could probably go build that today and figure out what that you know, talk to cloud code and figure out what to go build. Like, that is you know, it's it's a triple threat in this in this environment right now. What are you seeing at the enterprise level? Cuz they're not even close to there. So, if you're at a big enterprise, if you're at a Pepsi or something like that, you're still trying to take on the best practices. What realistically, what can they take on and how do they take them on? Yeah, I I mean I I think what I'm seeing in enterprises is like they're still innovating at um you know, there's in a you know, I I don't want to say that there's no innovation happening there at all. Like right now, there is yeah, all these teams are all using the coding agents and I think feeling the unlock of those those tools in their own day-to-day workflows. Um and so I I think what I'm seeing coming out of the teams right now, even there is like one amazing products that use AI to make the experience of that product useful. Um two, I think there's usually a massive uh especially larger companies, you have silos of data and people who you might have access to some information, other teams don't have access, and there's actually
(接上,Aakash 插话)99% 的人都不在 AI 原生公司里,所以他们不信我们。所以我得确认一下这是真的:你说的意思是,有时候来了一个问题,你的 PM 会判断它够重要,然后要么 PM 自己、要么工程师做出原型并打磨到可上生产的状态,然后你们当天就把它发出去?(Aparna:)是的,这就是真实在发生的事,各位。所以她自己都这么说了。那 PM 的角色又是什么呢?PM 是不是到这一步就得变成工程师了?(Aparna:)我觉得在这些 AI 原生团队里,我看到的是 PM 和工程师之间的差距已经分不清了。因为当写代码——
[1:06:06] Aparna
a really great piece that um Jaya Gupta some you should follow on Twitter, uh kind of shared a couple weeks ago now that's gone super viral around context graphs. Um and what a context graph is is essentially can you give your agent access to the agents are only as good as how much context that they actually have. Um and then of course the harness that's built on top of that, you know, has access to that context. And so instead of all that information and data being in completely different silos and and people operating in these silos, can you give one unlock for agents is that can you give it access to the context from different environments? And what that does is it actually makes people kind of kind of bridge the gaps across across different teams in ways that probably weren't possible before. And so figuring out how agents consume the context within an organization is going to be probably one of the biggest problems Um I mean it's probably one of the biggest unlocks challenges and unlocks that we're going to see this year. So if you're a product leader at one of the enterprise companies you're seeing what you just demoed for us you're saying okay how can I bring my company towards that what's sort of the step-by-step road map I should be implementing over the next say 12 24 months. Well first I think as an individual IC I think build like building and what I just shared right now of like you'll read a lot of just stuff on AI Twitter of everyone kind of um you know everyone kind of sharing every latest new model and every latest new tool out there I think what I would just highly recommend for any IPM is start by building start by building very simple like this example that we just did today it doesn't even need to be an external facing agent that you
——变得容易这么多以后,其实你知道,这又回到了我们今天播客一开始聊的那个点:今天的 alpha(超额优势)就是产品品味。所以那些懂产品品味、懂客户想要什么、懂怎么交付一个真正惊艳体验的人,速度会快得离谱。所以那种能从「这是痛点,这是我认为非常惊艳的体验」一路走下来、还是个三料全能的 PM——他们会说「我今天大概就能把它搭出来,去跟 Claude code 聊聊该搭什么」——那在当下这种环境里就是三料全能。(Aakash:)那你在企业级(大公司)这边看到的是什么情况?因为他们离那个状态还差得远。如果你在一家大企业,比如百事这种公司,你还在努力吸收各种最佳实践。现实点说,他们能落地哪些东西、又该怎么落地?(Aparna:)是的,我是说,我觉得我在企业里看到的是,他们仍然在创新——我不想说那里完全没有创新发生。现在确实是,所有这些团队都在用编码 agent,而且我觉得他们也在自己日常工作流里感受到了这些工具带来的解锁。所以我现在看到这些团队里冒出来的,第一是用 AI 让产品体验真正有用的那些惊艳产品;第二,我觉得通常会有一个巨大的——尤其是大公司——你会有一个个数据孤岛,有些人能拿到某些信息、别的团队却拿不到,而其实——
[1:08:08] Aparna
need to publish can it just be an internal kind of tool that you use to actually help you unlock make one big unlock today. Um that's huge that's huge because think about you know if this tool that we just vibe coded in an hour now I'm going to go use it to figure out okay well what are my top pains like what are this like you can imagine the next step after that is well can I get an agent to actually go and put up a draft PR for one of these can I get an agent to actually then review that PR and do the code review on that and then you know eat like the process to go from identifying a pain point to solve and then releasing that could have taken months in the past can now that entire thing can be shortened to a span of like we were saying a day. And if you just started with like if that could be your day. Um and what does everything need to look like in order to deliver on that? Um I think it changes the game for individual ICs. So first I'd say start by start by building. Um it's the most in a biggest unlock to um I as you're building it's kind of important to figure out what are the systems that you need in place in order for you to it's easy to kind of build something and then say oh it doesn't work. Like I'm just going to you know how many times has happened to you where you're like it's not working like I'll just kind of scratch that idea and kind of let it sit. Um I think the the most curious and the the most kind of you know uh curious of the PMs are typically this is where having a data layer like Arize uh and the observability platforms are really helpful is that you know you might not know like why your agent gave you a bad response or why the outcome wasn't that great what it was doing. And so getting observability to understand kind
——有一篇非常棒的文章,是 Jaya Gupta——你应该在 Twitter 上关注一下她——大概几周前发的,现在已经传疯了,讲的是 context graph(上下文图谱)。所谓 context graph,本质上就是:你能不能让你的 agent 拿到——agent 的好坏完全取决于它实际拥有多少 context。然后当然还有搭在它上面的那层 harness(执行框架),能访问到这些 context。所以与其让所有这些信息和数据散落在完全不同的孤岛里、大家各自在孤岛里工作,你能不能给 agent 一个解锁点:让它能跨不同环境去拿到 context?这么做的效果是,它真的能让人以前不可能的方式跨不同团队去弥合鸿沟。所以搞清楚 agent 怎么在一个组织内部消费 context,很可能会是今年我们会看到的最大问题之一——我是说,很可能是最大的挑战和解锁点之一。(Aakash:)那如果你是某家企业公司的产品负责人,你看到了你刚给我们演示的东西,你会说「好,我怎么把我公司带到那个状态?」在接下来比如说 12 到 24 个月里,我该一步步实施怎样的路线图?(Aparna:)嗯,首先我觉得作为一个个人 IC(独立贡献者),我觉得就去搭——就像我刚刚分享的那样。你会在 AI Twitter 上读到一大堆东西,每个人都在分享每一个最新的模型、每一个最新的工具。我特别想强烈建议任何一个 PM:先从动手搭开始,从搭非常简单的东西开始,就像我们今天做的这个例子。它甚至不需要是一个对外发布的 agent,你也不需要——
[1:10:13] Aparna
of like we were talking about with just a simple example of like the tennis players. Like how do they look at their plays and figure out what went wrong and how do they do that 1% better every single day? If you could n plus one your output every single day I I think that the story is no longer about observability oh looking at your data the story's about self-improvement and improvement of yourself as a PM but also improvement for the products that you are building. So we used Arize's open source sort of Phoenix platform and then we used Arize the paid platform to do this. Those are two options. How does somebody make a decision? What does the overall ecosystem look like and why would they choose a rise? Yeah, great question. Um so arise Phoenix, which was kind of the open source one that we pulled all the GitHub issues from today, um is an amazing option if you cannot send your data to an external platform. And for most enterprises, most teams building any agents that have any PII data, it's just a reality is that they want to self-host um some initial observability so that they can get a feel and get started and get an unlock. And so arise Phoenix is the I think even Hamel's tweeted this before, his most favorite open source um tool for observability is arise Phoenix. It's got super permissive uh license. You know, it's got almost everything that you just saw in the demo today out of the box for you. Um and all the skills that I shared kind of using cloud code, there's all those skills exist for Phoenix, too. So you can just go, open up, build an agent, and say, "Hey, help me instrument it. Help me figure out insights for my traces. Help me go write evals." Phoenix will actually go and do all of that for you today. Um typically where teams start to feel
——发布它。它能不能就是一个你自己用的内部小工具,真正帮你今天就拿下一个大的解锁?这太关键了,太关键了。因为你想想,如果像我们刚刚一小时里 vibe coding 搭出来的这个工具,现在我拿去用它来搞清楚「好,我最头疼的痛点有哪些」——你可以想象再往下一步:我能不能让一个 agent 真的去帮其中一个痛点提一个草稿 PR?我能不能再让一个 agent 去 review 那个 PR、做代码评审?然后,从识别出一个要解决的痛点、到把它发布出去,这整个流程在过去可能要花好几个月,现在整件事可以缩短到——像我们刚说的——一天之内。如果你就从这里开始,如果那能成为你的一天,那为了交付这个,所有东西又得长成什么样?我觉得这对个人 IC 来说是改变游戏规则的。所以首先我会说,从动手搭开始。这是最大的解锁。然后,在你搭的过程中,搞清楚你需要哪些系统就位很重要,因为很容易出现这种情况:你搭了个东西,然后说「哦,它不行」。你有多少次是这样的——你觉得「它不工作,那我就把这个想法划掉、先搁着」。我觉得最好奇的那些 PM,通常这正是 Arize 这样的数据层和 observability 平台真正能帮上忙的地方:你可能不知道为什么你的 agent 给了你一个糟糕的回答,或者为什么结果没那么好、它当时到底在干什么。所以拿到 observability 去理解——
[1:12:16] Aparna
you know, the the paid kind of platform uh the enterprise platform kind of makes sense is when obviously data volume starts to scale. We have teams that send us you know, just the volume of I I think it's a good thing is that these agents are starting to find product market fit in this environment right now. It's that LLMs uh the models are you know, getting better. Products are starting to find product market fit. And so we're starting to see, you know, almost like terabytes of data. Um and so it is um the volume and the scale is a big reason why um, you know, for teams that that need that as our agents start to get mature, uh, it makes a ton of sense for you to kind of have a more scaled out platform for observability. This is where Arize AI x is kind of uniquely fit to solve that problem. Um, we we do this really well because we've actually had to invest in our own data store that we've been building for a while now, ADB. Um, and it's a data store that's designed for AI workloads from day one. So, let's say I'm figured out I need to pay for it. I have the huge amount of data. How do I decide who to work with? The reason to pick Arize is really where the open and independent most independent platform out there. We are independent of framework. We don't actually care what framework you use. Um, we have teams using, you know, everything from LangChain the Cloud Agents SDK to teams that are building without a framework. Um, and so we we deeply don't we're agnostic of whatever framework you you use. Um, the second thing is we we deeply believe in the independence of your data. All of our trace data that we collect lives in open formats. Uh, you can actually using our ADB data fabric, um, that data can be directly sent back to your data warehouse. And
——就像我们刚用网球运动员那个简单例子聊的:他们怎么去看自己的每一拍、搞清楚哪里出了错,又怎么做到每天都比昨天好 1%?如果你能让自己每天的产出都 n+1,我觉得这个故事就不再只是关于 observability、关于看你的数据了,这个故事是关于自我提升——既是作为 PM 提升你自己,也是提升你正在搭的那些产品。(Aakash:)所以我们用了 Arize 开源的 Phoenix 平台,然后我们又用了 Arize 的付费平台来做这件事。这是两个选项。那一个人该怎么做选择?整个生态长什么样?他们为什么会选 Arize?(Aparna:)问得好。Arize Phoenix,也就是我们今天用来把所有 GitHub issues 拉出来的那个开源版本,如果你没法把数据发到外部平台,它是一个非常棒的选择。对大多数企业、大多数搭任何带 PII(个人身份信息)数据 agent 的团队来说,现实就是他们想自托管一套初步的 observability,这样能先上手感受一下、先跑起来、先拿到一个解锁。所以 Arize Phoenix 是——我觉得连 Hamel 之前都发推说过,他最喜欢的 observability 开源工具就是 Arize Phoenix。它的 license 极其宽松,你今天在演示里看到的几乎所有东西都开箱即用。而且我刚分享的那些用 Claude code 的 skills,所有那些 skills 在 Phoenix 上也都有。所以你可以直接打开、搭一个 agent,然后说「嘿,帮我做埋点(instrument),帮我从 traces 里找出洞见,帮我写 evals」。Phoenix 今天就真的能帮你把这些全做了。一般来说,团队开始感觉到——
[1:14:16] Aparna
the reason that's really powerful is because you don't want our you don't want your agent trace data, which is so valuable, to be locked inside of a proprietary platform. We make it accessible so that you can actually use the agent trace data as uh, part of your context graph. Um, we're also independent of instrumentation. I If you don't know, we're actually the inventors of open inference. Our competitors, every single one of them that you mentioned, all use our instrumentation and they've actually linked to it in their docs. And so, we actually own um kind of we've built probably the richest telemetry um and it kind of shows in the fact that our instrumentation is widely adopted in the ecosystem. Um and then the the last one I think is just I think we've been consistently one of the most innovative in the market. We were actually the first to shipping LLM as a judge. If you go back to 2023, you'll look at Phoenix in the repo and you'll see kind of LLM as a judge. We were the first to uh release open inference instrumentation, Alex that you saw in kind of the product. We were the first to actually have an agent built into our product. The skills that you're actually looking at that we kind of showed how you use, all of those skills um we were actually the first to have and release them. Um Hamel actually did uh a talk with Mikio on this about it, our open-source lead. Um and then I was I was mentioning we have kind of the first and only way right now in market to actually take all of these agent traces and have them as standard formats as part of your context graph. And I think it just shows, you know, we are we're probably uh the fastest innovator in the space right now. What are What are the things, if somebody has just 2 hours this weekend, that they should
——付费平台、企业版平台开始变得划算,显然是在数据量开始往上涨的时候。我们有些团队发给我们的——我觉得这其实是件好事——就是这些 agent 在当下这个环境里开始找到 product market fit 了。LLM、模型在变好,产品开始找到 product market fit,所以我们开始看到几乎是 TB 级别的数据量。所以这个量级和规模就是一个很大的原因,对那些有这种需求的团队——随着我们的 agent 开始走向成熟,搭一个更能横向扩展的 observability 平台对你来说就非常合理了。这正是 Arize AI 独特擅长解决的问题。我们做得特别好,因为我们其实不得不投入去搭自己的数据存储,已经搭了挺长一段时间了,叫 ADB。它是一个从第一天起就为 AI 工作负载设计的数据存储。(Aakash:)那好,假设我已经想清楚我需要付费了,我有海量数据,那我怎么决定该跟谁合作?(Aparna:)选 Arize 的理由,真的在于我们是市面上最开放、最独立——最独立的平台。我们独立于框架,我们其实不在乎你用什么框架。我们的团队从 LangChain、Claude Agents SDK 到不用任何框架自己搭的都有,所以我们对你用什么框架是真的无所谓、保持中立的。第二点是,我们深信你的数据应该是独立的。我们采集的所有 trace 数据都以开放格式存储。你其实可以用我们的 ADB 数据底座(data fabric),把那些数据直接发回你自己的数据仓库。而——
[1:16:08] Aparna
concretely go do and take away besides just they've watched this episode, but now they're going to actually make impact in their career? If you uh if you have any 2 hours uh this weekend, I would say literally what we just did right now, which is build build an agent for yourself. Whatever would take away a couple hours of your week every week. Like just something repetitive that you do every single week. And by the way, this isn't just for like if you're someone who's in product marketing and you're writing release notes every week. Like, what is just a workflow that you do every single week that takes a couple hours of your week, try to build an agent to go do that. And I think what you'll learn out of that is one, how insanely easy it is with Claude code, um and then you'll also on the other hand realize how you know, how much work it takes to actually make it really good. And so, to make it get better, um past that initial kind of vibe code, the evals and the observability are so important. And so, you know, I said this in the beginning, but any product person that has used observability and is looking at their traces and looking at your evals, you're probably already in the top 1% of PMs in the world right now. What are the biggest mistakes PMs are making when they do evals? I think the biggest ones is um not first not starting with actual trace data. Um I think if you're just starting with uh kind of what you think are problems, that's really hard. Um like even the skills, for example, that we used today, um that that Claude was using to build the evals, what's powerful about it is that it's actually we're trying to instill best practices. It's actually looking at all of the trace data to help and suggest what the right evals could be. Um so, I
——这之所以特别有威力,是因为你不会想让你的 agent trace 数据——它这么有价值——被锁死在一个专有平台里。我们让它可访问,这样你就能真的把 agent trace 数据当作你 context graph 的一部分来用。我们在埋点上也是独立的。如果你不知道的话,我们其实是 OpenInference 的发明者。我们的竞争对手,你提到的每一家,全都在用我们的埋点,而且他们其实在自家文档里都链接到了它。所以我们其实拥有——可以说我们搭出了大概是最丰富的遥测(telemetry),这一点也体现在我们的埋点在整个生态里被广泛采用上。然后最后一点,我觉得就是我们一直是市场上最具创新力的玩家之一。LLM-as-a-judge(用 LLM 当评判)就是我们第一个做出来的。你回到 2023 年去看 Phoenix 的 repo,就会看到 LLM-as-a-judge。OpenInference 埋点——也就是你在产品里看到的那个 Alex——是我们第一个发布的。我们也是第一个真正把 agent 内建进自己产品里的。你现在看到的、我们演示过怎么用的那些 skills,全都是我们第一个做出来并发布的。Hamel 其实还跟 Mikio——我们的开源负责人——专门做过一期聊这个的访谈。然后我刚提到的,我们现在还有市面上第一个、也是唯一一个能真正把所有这些 agent traces 以标准格式纳入你 context graph 的方式。我觉得这恰恰说明,我们大概是这个领域里当下迭代创新最快的玩家。(Aakash:)那如果有人这个周末只有 2 小时,他们具体该去做点什么、带走点什么——除了「看了这一集」之外,还能真正——
[1:18:07] Aakash
I think PMs need to look at the evals don't just come out of magic, they come out of your they come out of traces. All right, everybody. I'm going to put up a Rise's pricing page for you. This is how much Rise costs. Now, here's the cool thing. If you want to get AX Pro for 12 months for free for your team because you're convinced you want to create self-improving agents, you can do that with Akash's bundle. Or you can just use the free options that she's talked about right now, Phoenix and AXE Body Spray, to get started. It's that simple. I highly recommend every AIPM master the AIEval skill. Arise is one of the easiest ways to do it. Aparna, thank you so much for lending your expertise. Awesome. Thank you so much, Akash, and uh it was awesome to be here. I hope you enjoyed that episode. Couple things you can do to support the show. One, comment. Two, review. Those ratings and reviews really help other people understand the value and the production that we are putting into this, right? This wasn't an easy episode to produce. We put in a ton of pre-work. We edited it for you. We brought in the best guests. If you don't mind sharing a rating and review, sharing the episode with others, making sure you are subscribed, that really helps the show do bigger and better productions. I'll see you in the next episode. Here is one of those that YouTube thinks would be a great fit for you.
我觉得 PM 得去看这些 evals——它们不是凭空变出来的,是从你的 trace 里来的。好了各位,我给大家放一下 Arize 的定价页面,这就是 Arize 的收费情况。重点来了:如果你已经被说服、想去打造自我改进的 agent,想免费拿到 12 个月的 Arize 团队版(Pro),可以走 Aakash 的捆绑套餐。当然,你也可以先用她刚才提到的免费方案——Phoenix——上手,就这么简单。我强烈建议每一位 AI PM 都把 AI eval 这项技能练到家,而 Arize 是最省事的入门方式之一。Aparna,太感谢你来分享你的专业见解了。很棒。也特别感谢你,Aakash,能来这儿真的太棒了。希望你喜欢这一期。有几件事你可以做来支持这档节目:第一,留言评论;第二,给个评分。这些评分和评价真的能帮到其他人,让他们了解我们在这上面投入的价值和制作用心,对吧?这一期可不好做,我们做了大量的前期准备,给你精心剪辑,请来最好的嘉宾。如果你不介意的话,给个评分和评价、把这期节目分享给别人、确保你已经订阅,这些都能帮节目把制作做得更大更好。我们下期见。这里有一个 YouTube 觉得很适合你的视频。