How to Build AI Evals Step-by-Step | Daniel McKinnon | Product Growth
频道: Aakash Gupta
视频: https://www.youtube.com/watch?v=ztN6bE_FuQQ
原文语言: en
统计: 共 69 轮 · Daniel 47 · Aakash 22
[0:00] Daniel
The Metas and the Googles and all the other large companies have to reinvent themselves right now in the age of AI. Every single PM is going to start building AI features. You cannot exist as a PM without understanding this.
Meta、Google 这些大公司,在 AI 时代都必须重新发明自己。以后每一个 PM 都会开始做 AI 功能。不懂这套东西,你根本没法继续当 PM。
[0:13] Aakash
Meet Daniel McKinnon, former PM on the llama models at Meta, [music] now a startup founder and a master of eval.
这位是 Daniel McKinnon,曾在 Meta 负责 Llama 系列模型的 PM,现在是创业公司创始人,也是 eval(模型效果评测)方面的高手。
[0:20] Daniel
PMs right now from a career perspective are in a really tough situation. The average PM is an orchestrator, a motivator, and an analyst. But a lot of this is easy to do with AI.
从职业发展的角度看,PM 现在的处境其实挺难的。普通 PM 干的活无非三样:协调资源、鼓舞团队、做分析。但这几样,AI 做起来都不难。
[0:30] Aakash
What really is the difference between product management at Meta versus Google?
在 Meta 做产品和在 Google 做产品,真正的区别到底在哪?
[0:33] Daniel
Meta is like a much much more aggressive culture. Uh in many ways, Google is considered to be more of an engineeringled company whereas Meta is more of a productled company.
Meta 的文化要激进得多。很多方面来说,Google 被认为是一家工程驱动的公司,而 Meta 更像是产品驱动的公司。
[0:42] Aakash
If you're a PM who's [music] never worked on an AI feature before, when would you be going through this ebalance process?
如果一个 PM 之前从没做过 AI 功能,那他大概会在什么阶段才需要走这套 eval 流程?
[0:48] Daniel
You work at Pinterest and you want to have uh better image generation that doesn't look like sloth that actually pleases users. You need to think about from day zero. What does success look like?
假设你在 Pinterest 上班,想把图片生成做得更好,不要生成那种一眼假的 AI 垃圾图,而是真能让用户满意的图。那你从第 0 天就得开始想:怎样才算成功?
[0:58] Aakash
If it's the best way to measure it, we've got to learn it. Where should we start?
如果这确实是最好的衡量方式,那我们就得学会它。该从哪儿开始?
[1:02] Aakash
Let me walk through. Before we get into today's show, please take a second to check that you're subscribed on YouTube and following on Apple and Spotify podcasts. If you want access [music] to all of my favorite AI tools, I've gotten them to give you an entire year of their paid plans. Check out bundle.ac. akashg.com for an entire year of bolt new air table speechify descript magic patterns linear dovetail arise and mobin and now into today's show Daniel welcome to the podcast
我来一步步讲。进入今天的正题之前,麻烦花一秒钟确认一下你已经在 YouTube 订阅了本频道,并在 Apple Podcasts 和 Spotify 上关注了本播客。如果你想用上我最喜欢的那些 AI 工具,我已经帮你谈下了整整一年的付费版权限。去 bundle.aakashg.com 看看,Bolt.new、Airtable、Speechify、Descript、Magic Patterns、Linear、Dovetail、Arise、Mobbin 全都送你一年。好,进入今天的正题。Daniel,欢迎来到播客。
[1:39] Daniel
yeah thanks so much for having me and uh it'll be fun to talk about this stuff
谢谢邀请,能聊这个话题挺有意思的。
[1:43] Aakash
so I want to start with your article you had this provocative claim and this funny meme here do evils replace the PRD what is the role of evils
我想先从你那篇文章聊起。你在里面抛了一个挺挑衅的观点,还配了张搞笑的梗图——「eval 是不是取代了 PRD?」eval 到底扮演什么角色?
[1:52] Daniel
yeah so I was like a little bit spicy in saying it replaces the PRD because without a product strategy or a particular customer like your product is nothing but again that's a paragraph or even a sentence depending on what the product is. The majority of most PRDs that I've seen in my career have spent most of the document talking about specifics, how it will work, how it will behave in certain situations, how the user can expect to get value from that. But that gets really turned on its head in this kind of Gen AI world where these products really need to do like everything or at least a lot more things than previous products. And it's very hard to describe that, say, oh, this thing just does everything. And the best way to actually communicate what the product should do is through examples. And that's what an eval is. It's really just like a trivia question for the model. And it's saying this is like the shape of the things the model needs to do well. And if the model does it well, it means it's getting these answers. And if it gets these answers, the users will probably like it. And if the users probably like it, let's ship it into prod and see if they actually like it with an online eval. But the key way to communicate how a product should work in this kind of Gen AI era is is it performing well on an offline eval? And if not, you either need to change the model, change the harness, or change the product. It's possible that what you want to do is not possible with the models today. But it's better to find that out early with an offline eval versus just shipping it to prod and getting frustrated users.
对,说「eval 取代 PRD」确实有点故意刺激人。因为没有产品策略、没有明确的目标用户,你的产品就什么都不是。但话说回来,策略这部分往往一段话就写完了,有的产品甚至一句话就够。而我职业生涯里见过的绝大多数 PRD,篇幅主要都花在细节上:功能怎么运作、在某某情况下会怎么表现、用户能从中获得什么价值。可到了生成式 AI 这个世界,这套写法完全被颠覆了——这类产品几乎什么都要能干,至少要干的事比过去的产品多得多。你很难把这个描述清楚,总不能写一句「这玩意儿啥都能干」。真正把产品该做成什么样讲明白的最好方式,是给例子。而 eval 就是这个例子。它其实就是给模型出的一道道知识问答题,等于在说:模型要做好的事情,大致就长这个样子。如果模型答得好,说明它给出的是这些答案;如果它给的是这些答案,用户大概率会喜欢;用户大概率会喜欢,那就上生产环境,再用线上 eval 看看他们是不是真的喜欢。但在生成式 AI 时代,说明一个产品该怎么运作,关键就一句话:它在线下 eval 上表现好不好?如果不好,那你要么换模型,要么改 harness(模型外面那层调度和工具框架),要么改产品。也有可能你想做的事,用今天的模型根本做不出来。但这件事早点靠线下 eval 发现,总好过直接上线,然后收获一堆被激怒的用户。
[3:22] Aakash
And for people who don't quite understand that nuance, what's an offline eval versus just shipping to prod and looking at them?
对不太理解这个区别的人来说,什么叫离线 eval(offline eval,线下评测),跟直接上线到生产环境去看效果又有什么不一样?
[3:28] Daniel
Yeah. Yeah, this is a really really important nuance and I touched on it in this blog post, but when we usually talk about evals in this AI world is something that's run offline, there's like a little bit of gray areas in terms of RL environments and stuff, but think about it as like a pre-baked set of trivia questions that you ask the model. So, for example, let's say you have a recipes website and you want to tell users how to make their favorite kinds of ice cream. An offline eval would be a prompt set, say 100 prompts of different ice creams that users might like. And the answer key would be uh a correct answer or a plausibly correct answer with a way to score whether it's good. So that's run offline during the development of your product and you use that as a proxy for real user traffic. Once you do well on your offline evals, you can ship that product online and you have a website and it lets users generate ice cream recipes and you think because you are a good PM and you really thought deeply about what the customer wanted that performance on that offline email set will reflect the satisfaction of the user in the online eval. Um, this obviously doesn't always happen, but that's the idea. And it's the best way to measure whether a Genaii product is likely to satisfy users.
对,这个区别特别特别重要,我在那篇博客里也提到过。我们在 AI 圈子里说 eval 的时候,通常指的是离线跑的那种——当然像 RL environment(强化学习环境)这类东西中间有些灰色地带——但你可以把它理解成一套事先准备好的「问答题库」,拿去考模型。举个例子,假设你做一个菜谱网站,想告诉用户怎么做他们最爱的冰淇淋。那离线 eval 就是一组 prompt,比如 100 条不同冰淇淋的提问。而「标准答案」就是一个正确答案,或者一个说得过去的正确答案,外加一套判分方法。这套东西在你开发产品的过程中离线跑,你把它当成真实用户流量的代理指标。等离线 eval 上的成绩不错了,你就可以把产品上线,做成一个网站让用户生成冰淇淋菜谱。你会指望——因为你是个好 PM,对用户到底想要什么想得足够深——离线那套题上的表现能反映出线上真实用户的满意度。当然这事儿不是每次都成立,但思路就是这样。而且这是判断一个 GenAI 产品能不能让用户满意的最好办法。
[4:48] Aakash
If it's the best way to measure it, we've got to learn it. Where should we start?
既然它是最好的衡量办法,那我们就得学会它。该从哪儿开始?
[4:51] Daniel
Let me walk through what's changed since I wrote that article. So the key thesis has remained true and this has been now two years in that an eval is the best way to communicate what your product should be doing and to explain to the engineering team working on making the product work what success looks like. This can be very very challenging in an AI world when they these products do so many different things that it's hard to necessarily understand what good is. But things have gotten a lot more complex. So back in the day, evals were really simple. My blog post basically just covered these simple evals. They were question and answer. And this seems crazy to think back 2 years about what these models looked like and what the use cases were like, but this is before cloud code and agentic coding and all of these crazy business applications that are getting built right now. uh claude co-work and the like. Really the core thesis two years ago was that Genai was essentially a search replacement. I don't know if everyone remembers when Google's stock tanked because uh this was going to be the replacement for search and fundamentally these were like question and answer products. So what I have right now is the benchmarks that OpenAI reported on GPT4. And you might remember a lot of these MMLU, Hela, Swag, ARC, Window, human eval human eval is ironically an automated eval of Python. Uh drop and these are all just question and answers. So I dropped one example here from MMLU. If you know the actual brightness of an object and its apparent brightness from your location, then with no information, you can
我先讲讲自从我写那篇文章以来有什么变了。核心论点到现在两年了还是成立的:eval 是传达「你的产品到底该做什么」的最好方式,也是向做这个产品的工程团队解释「什么叫做成了」的最好方式。在 AI 这个世界里这件事可能非常非常难,因为这些产品能干的事太多了,你很难说清楚「好」到底长什么样。但事情也确实复杂了很多。当年 eval 特别简单,我那篇博客基本就只覆盖了这种简单 eval——就是一问一答。现在回头想两年前的模型长什么样、用例是什么样,感觉挺不可思议的,那还是在 Claude Code、agentic coding(智能体写代码),以及现在正在被造出来的这一堆疯狂的商业应用——比如 Claude Cowork 之类——出现之前。两年前真正的核心论调是:GenAI 本质上是搜索的替代品。不知道大家还记不记得 Google 股价暴跌那会儿,因为这玩意儿被认为要取代搜索了,而这些产品本质上就是问答型产品。我这儿现在放的是 OpenAI 当年为 GPT-4 汇报的那些 benchmark(基准测试)。你可能还记得不少:MMLU、HellaSwag、ARC、WinoGrande、HumanEval——HumanEval 讽刺的是,它其实是一套自动化的 Python eval——还有 DROP。这些全都是一问一答。我这儿从 MMLU 里摘了一道题:如果你已知一个物体的实际亮度和它在你所处位置的视亮度,那么在没有其他信息的情况下,你可以估算出——
[6:35] Daniel
estimate a speed relative to you, b composition, c size, d distance from you. And if you want to go ahead and look at some of these just for like historical fun, uh they're they're all on hugging face. So this was actually a very very straightforward thing to do is you had to think about the types of questions users would ask and the types of answers they would expect and how to score that. So the key thing was just matching the questions to the domain of interest and scoring the answers. So if we go back into my post here, we can see a little bit how I said to do this. And uh first thing is just figure out your problem and it doesn't need to be perfect. What is a set of problems that your users might ask about? So the example I used was uh for example generating recipes from videos. I guess I have food on my mind because I randomly came up with the ice cream uh example earlier. And new problem is I have a video on my social media site and I want to be able to generate a recipe that somebody can use to make that thing as they're watching the video. And you say, "Okay, this is really well defined." Um, and now you have to measure h how do you know if it's good?
——A 它相对你的速度,B 它的成分,C 它的大小,D 它离你的距离。如果你想去翻翻这些题当个历史趣味看看,它们都在 Hugging Face 上。所以这件事当年其实非常非常直白:你得想清楚用户会问哪类问题、期待什么样的答案,以及怎么给答案打分。关键就是把题目对准你关心的领域,再把答案判分做出来。回到我那篇文章里,能看到我当时是怎么建议的。第一步就是先想清楚你的问题,而且不需要一开始就完美——你的用户可能会问的一组问题是什么?我当时举的例子是「从视频里生成菜谱」。看来我脑子里老是吃的,刚才随口就冒出个冰淇淋的例子。这个新问题是:我的社交媒体网站上有一段视频,我希望能生成一份菜谱,让人一边看视频一边照着做出那道菜。你会说:好,这定义得很清楚。接下来就得衡量了——你怎么知道它做得好不好?
[7:45] Daniel
You know, it might be formatted right. It might have all the ingredients listed. It might be written in the right style. And then you select all of these components and figure out how to judge if it's correct or hypothesize how to judge this correct. This can be a auto score. This can be another LLM. This can be a human. You can judge correctness many different ways. Once you have that, it's really a mechanical process to actually write the eval. Just come up with probably 100 prompts that are in this distribution. Can be less, can be more, but this is a typical size of an eval in a genai world. and just send it through the model and figure out what is a hard prompt, what is an easy prompt, and have something that scores maybe like 50%. Cuz you have to have room to run. If you create a very easy eval that scores 100%, there's no way for your engineering team to optimize on that. And if you create a very hard eval that scores 0%, you also don't even know if this is kind of possible with today's technologies. Um, and that's pretty much it. And once you have this uh you can give it to the team, they can improve on it. You can ship it out to users. If you achieve a score high enough, you can see if you're actually online performance matches what you expect it to do offline. And then uh this uh blog has a bunch of examples of like, you know, some kind of other nice tips. But why is this not that relevant today or why does it need to change today? The real answer is that models have generally saturated QA. This isn't 100% true, but when we think of a good model right now, we don't think of one that can answer a
可能是格式对不对,可能是配料有没有列全,可能是文风写得对不对。然后你把这些维度都挑出来,再想清楚怎么判定它是对的,或者先假设一套判定「对不对」的办法。这个判分可以是自动打分,可以是另一个 LLM(LLM-as-judge,用大模型当裁判),也可以是人来判。判正确性的方式有很多种。这些定好之后,真正写 eval 就是个机械活了:凑大概 100 条落在这个分布里的 prompt——可以更少也可以更多,但在 GenAI 这个世界里,这是一套 eval 的典型规模——然后丢给模型跑一遍,看哪条难、哪条容易,最后凑出一套大概能得 50 分的题。因为你得留出提升空间。你要是做了一套特别简单、直接考 100 分的 eval,那工程团队就没法拿它去优化了;你要是做了一套特别难、考 0 分的 eval,你连这事儿在今天的技术条件下到底可不可能做到都不知道。基本就是这样。有了这套东西之后,你就可以交给团队,他们照着往上提;也可以发布给用户,如果分数够高了,你就能看看线上的实际表现是不是跟你在离线时的预期对得上。那篇博客里还有一堆例子和其他一些不错的小技巧。但为什么这套东西今天不那么适用了、为什么必须变?真正的原因是:模型在 QA(问答)上基本已经饱和了。这话不是百分百准确,但我们今天说一个模型好,想的已经不是它能不能答对——
[9:18] Daniel
relatively challenging high school physics question like the uh example I gave above. We think of models that are like winning gold medals at international math Olympiads. Like there's almost no question that any human beyond some super super specialist can ask the model to do and it not have a good answer back. And also QA is not the most useful application now. I mean I think we all remember this narrative that chat GPT was going to be the next great consumer app and they were going to get all these users and it's QA and you're answering all these problems and you know Google did a generative search experience. But if you look at all the headlines in Genai right now, it's not QA, it's agents. And this is actually reflected with how the labs communicate progress. Here's a quick word from our sponsors. If you're building anything that uses live data from the web, eventually you hit the same wall, an agent, a research tool, trends, dashboard. They all need fresh data.
——一道我上面举的那种偏难的高中物理题了。我们想的是能在国际数学奥赛上拿金牌的模型。基本上除了极少数超级专才,普通人能问出来的问题,模型都能给个像样的回答。而且 QA 现在也不再是最有用的应用了。我想大家都还记得那套叙事:ChatGPT 会成为下一个伟大的消费级应用,会吸走所有用户,而它做的就是问答,回答你各种问题,Google 也搞了个生成式搜索体验。但你看现在 GenAI 领域所有的头条,主角已经不是 QA 了,是 agent(智能体)。这一点在各家实验室怎么汇报进展上也体现得很明显。(插播赞助)先来听一段赞助商的话。如果你在做任何需要网上实时数据的东西——一个 agent、一个研究工具、一个趋势看板——早晚都会撞上同一堵墙:都得要新鲜数据。
[10:15] Aakash
Scraping that data is the worst part. Captas, proxy, layouts that change every week. It's a whole side project you didn't sign up for. That's where SER API comes in. SER API gives you clean structured results from Google, YouTube, Bing, Google News, Google Scholar, and more. One API call, one clean JSON response. They handle the captions, proxies, and layout changes for you. Take the Google Scholar API as one example. [music] Say you're building a research assistant or pulling sources for a literature review. You hit one [music] endpoint, you get peer reviewed articles back with full text or metadata, titles, links, publications, citation info, all of it across publishers and formats. No scraping a dozen publisher sites and gluing the data together yourself. The same idea extends across the rest of their APIs. Real-time Google search for an agent, pre-classified images for training data, Google News for monitoring, 99.9% uptime, [music] 1.2 second response time. Get started with 250 free credits. Link is in the description or scan the QR code on screen. Thanks to SER API for sponsoring. Are you looking to up your AI product management chops? I highly recommend [music] the AI product management certification by product faculty. It has a 47 with 1,249 reviews on Maven for a reason. I myself took the course back in 2024 and it was awesome.
而抓这些数据是最烦人的部分:验证码、代理、每周都在变的页面结构。这活儿本身就成了一个你根本没打算接的副业。SerpApi 就是干这个的。SerpApi 给你从 Google、YouTube、Bing、Google News、Google Scholar 等等地方拿到干净的结构化结果。一次 API 调用,一份干净的 JSON 响应。验证码、代理、页面改版这些他们全帮你扛了。就拿 Google Scholar API 举例:假设你在做一个研究助手,或者在为文献综述扒来源,你打一个接口,就能拿回同行评议的论文,带全文或元数据——标题、链接、刊物、引用信息,全都有,跨出版商、跨格式。不用自己去爬十几家出版社的网站再把数据拼起来。同样的思路贯穿他们其他的 API:给 agent 用的实时 Google 搜索、给训练数据用的预分类图片、用来做监控的 Google News,99.9% 可用性,1.2 秒响应。现在开始有 250 个免费额度,链接在简介里,或者扫屏幕上的二维码。感谢 SerpApi 赞助本期。另外,你想提升自己的 AI 产品管理功力吗?我强烈推荐 Product Faculty 的 AI 产品管理认证课。它在 Maven 上有 1249 条评价、4.7 分,是有原因的。我自己 2024 年上过这门课,非常棒。
[11:42] Aakash
Since then, they have upgraded it. So now you get to learn from product leaders at OpenAI and Enthropic. On top of that, Powell Hearn, author of the product compass, leads the build [music] labs. So you will go from theoretical knowledge about AIPM to a very practical course. [music] It's going to help you identify AI leverage opportunities. It's going to help you design trustworthy AI experiences. It's going to help you systematically optimize outputs for accuracy and relevance, build rigorous evaluation suites, architect AI agentic systems that work, and select the perfect LLM for [music] your use case. It's normally $2,500, but you get a discount when you use my link. The next cohort starts June 22nd and goes to August 9th. So, do check it out with my link in the description. They have been one of my longest sponsors for a reason. I trust this product and I think you should consider the cohort. So back here two years ago if you look at what how open AAI communicated progress it was these five eval I guess six evals if we fast forward to look at how anthropic communicated progress for opus 4.8 you can see they're using entirely different benchmarks you don't see any continuity it's partially because those eval are saturated and opus 4.8 8 would score effectively 100% on all of them. But it's partially because the task is really different. You'll notice we have agentic coding, agentic terminal coding, multidisciplinary reasoning. This is actually some agentic reasoning, agentic computer use knowledge work. This is actually agentic knowledge work and agentic financial analysis. All this means is what the core model task is is
从那以后他们又升级了课程,现在你能跟 OpenAI 和 Anthropic 的产品负责人学。除此之外,《The Product Compass》的作者 Paweł Huryn 来带实操 lab。所以你会从关于 AI PM 的理论知识,直接走到一门非常实战的课。它会帮你识别 AI 的杠杆机会、设计值得信赖的 AI 体验、系统性地优化输出的准确性和相关性、搭建严谨的评测套件(evaluation suite)、把能真正跑起来的 agentic 系统设计出来,以及为你的场景挑对 LLM。原价 2500 美元,用我的链接有折扣。下一期从 6 月 22 日开到 8 月 9 日,链接在简介里,去看看。他们是我合作时间最久的赞助商之一,这是有原因的——我信任这个产品,也觉得你该考虑一下这期课程。好,回到正题。两年前你看 OpenAI 是怎么汇报进展的,就是那五个 eval——算下来大概六个。快进到 Anthropic 怎么汇报 Opus 4.5 的进展,你会发现他们用的完全是另一批 benchmark,一点连续性都没有。一部分原因是那些老 eval 已经饱和了,Opus 4.5 在上面基本都是接近满分。但另一部分原因是任务本身真的变了。你会注意到这些名目:agentic coding(智能体写代码)、agentic terminal coding(智能体终端编程)、多学科推理——这其实也算某种 agentic 推理——agentic computer use(智能体操作电脑)、knowledge work(知识工作)——这其实是 agentic 知识工作——还有 agentic 财务分析。这一切说明的是,模型的核心任务已经不再是——
[13:23] Daniel
no longer to get a prompt from a user and come back with an answer. But it is actually to get a task from a user that requires many many steps. Some of these steps might involve just thinking which is called reasoning in this world. Some of it might involve tool calling something like search. Some of it might involve more advanced tool calling like something we'll go over today. And this is a totally new paradigm of writing evals because you're no longer thinking about QA. You're thinking about tasks. Fortunately for us, the framework is largely the same. We still have to define the problem. We still have to be good PMs and know what we're solving. We still have to collect representative prompts. I call this Goldilock style. Again, they can't be too hard and they can't be too easy. There has to be some room to run. A typical good eval will have something like 25% 50% success rate and then over you know months that will go to 100% and then you'll have to throw it away and create a new one that is harder. And then you also have to figure out how to score. And one things that has changed is QA is relatively easy to score with humans worst case scenario.
——接住用户一个 prompt 然后回一个答案了,而是接住用户一个需要很多很多步才能完成的任务。其中有些步骤可能只是「想」,在这个圈子里叫 reasoning(推理);有些步骤可能是调工具,比如搜索;有些可能是更高级的工具调用,像我们今天要过的那种。这是一套全新的写 eval 的范式,因为你想的不再是问答,而是任务。好在对我们来说,框架大体上还是那一套:还是得定义问题,还是得当个好 PM、清楚自己在解什么问题,还是得收集有代表性的 prompt。我把这一步叫「金发姑娘原则」(Goldilocks,不多不少刚刚好)——题不能太难,也不能太简单,得留出提升空间。一套好的 eval 典型的起点成功率大概在 25% 到 50%,然后过个几个月就会爬到 100%,那时候你就得把它扔了,再造一套更难的。再往下你还得想清楚怎么判分。有个变化是:QA 相对好判,最差情况还能上人工。
[14:32] Daniel
There's exceptions to this, of course. The reason these models are so bad at things like creative writing is because it's hard to score and there's different preferences and um different users like different things and why they're so good at math and coding is there is like some right answer and this is much easier to hill climb. But for to some extent QA style questions, you can ask human raiders to review worst case scenario if you can't find a better way to score it. With agentic work, it's much more challenging because the time horizon tends to be very long and the final output is a collection of many, many, many steps that it took. Some steps could be correct and lead to the wrong outcome. Some steps could not be correct. And it's just from a labor perspective and a um like defining success perspective, it's much much more important to get something that can be automatically scored uh to make more of these rollouts and understand how you can do more experiments. But again, it's largely the same. and the tasks are much longer time horizon. So kind of the goal of this podcast is to walk the audience through creating an actual agentic eval in real time. I want to caveat this.
当然也有例外。这些模型在创意写作这类事情上之所以那么弱,就是因为它难判分,人的偏好各不相同,不同用户喜欢的东西不一样;而它们在数学和代码上之所以那么强,是因为有明确的正确答案,这种爬坡(hill climb,指沿着指标往上优化)容易得多。但对 QA 类问题来说,实在找不到更好的判分办法时,你至少还能找人工标注员来评。到了 agentic 的活儿上,就难多了:任务的时间跨度往往很长,最终产出是很多很多步累积出来的结果。有些步骤是对的却导向了错误的结果,有些步骤本身就是错的。从人力成本的角度、从定义「成功」的角度,你都更需要一套能自动判分的东西,这样才能多跑几轮 rollout(一次完整的任务执行轨迹),才能理解怎么做更多实验。但话说回来,框架基本还是那一套,只是任务的时间跨度长得多。所以这期播客的目标,是带大家实时地做出一个真正的 agentic eval。我先打个预防针。
[15:42] Daniel
This is a little bit pre-baked. It is unrealistic in 45 minutes to come up with a brand new eval. This is kind of like a weeks or months problem of deep thinking, but we're going to kind of pretend and we'll we'll go through some of the steps together. So, first problem, I want to measure and improve the model's ability to help with clinical genomics. This is a problem that I care deeply about. It's one that I think uh can improve the world and it's something that I've launched a new startup to solve. And the problem is is that uh whole genome sequencing has become the absolute gold standard in diagnostics in NICU settings. So for sick babies, unfortunately interpreting the results of a whole genome sequence is very labor intensive and it limits access to this life-saving technology. So I wanted to see if I could distill some of this human expertise into a model to help broaden the accessibility of this technology. So first thing is like before we have to deeply understand the problem and I'm not going over these flowheets but this is just kind of how complex this is. Generally you start with the raw reads off the sequencer.
这个是有点提前准备过的。指望 45 分钟里从零做出一套全新的 eval 是不现实的,这基本是个要花几周甚至几个月深度思考的活儿。但我们就假装一下,一起把几个步骤走一遍。第一个问题:我想衡量并提升模型在临床基因组学上的帮助能力。这是我特别在意的一个问题,我认为它能让世界变好,我也为此创办了一家新公司。问题在于,全基因组测序(whole genome sequencing)已经成了 NICU(新生儿重症监护室)诊断里绝对的金标准。可对于这些重病的婴儿来说,解读一份全基因组测序结果非常耗人力,这就限制了这项救命技术的可及性。所以我想看看,能不能把这部分人类专业经验蒸馏进模型里,让这项技术能被更多人用上。那么第一步,在动手之前我们得深刻理解这个问题——我不会逐张过这些流程图,放出来只是想让你感受一下这有多复杂。一般来说,你是从测序仪出来的原始 reads(测序读段)开始的。
[17:03] Daniel
You do a lot of processing work to identify how the particular patient differs from the reference human genome and then you do another set of work to determine whether those changes to the genome matter. For example, if my genome were sequenced, I would get about a billion reads that are 150 base pairs long. they would come out in a giant text file and I would need to transform that into a diagnosis that says that this gene may or may not be responsible for this patient's condition. And there's a very structured way of doing this. So, I'm not going to spend too much time on this here because uh we're going to go over some real examples, but you really need to deeply understand the problem. This is why people like Anthropic and OpenAI are hiring investment bankers, accountants, lawyers. As you see job ads for all these vertical specific teams, you must deeply understand the problem. You will unlikely be successful in creating an eval for some topic if you don't have some background in it or haven't really educated yourself on it. So then let's say we've understood the problem. I think I understand this problem pretty well and by the end of this podcast you will too. Let's go to the prompts. So again we want to find this like Goldilocks set of prompts. So the first thing I like to do is just start with something easy. So you want to make sure the model can actually do this. And when I say the model for agentic stuff, I'm usually talking about the model plus the harness. So I'll use those words interchangeably. But this is a frontier model harnessed in a way that it can use these tools and it can do this reasoning
你得做大量处理工作,先搞清楚这位病人的基因组跟参考人类基因组(reference genome)差在哪儿,然后再做另一轮工作,判断这些差异到底有没有意义。举个例子,如果给我做测序,会得到大约十亿条 reads(测序读段),每条 150 个碱基对长,全都躺在一个巨大的文本文件里。我得把这一坨东西变成一句诊断结论:这个基因可能是、也可能不是造成这位病人症状的原因。这件事有一套非常结构化的做法。这里我不多展开,因为等下会过真实的例子,但关键是——你必须对这个问题有很深的理解。这也是为什么 Anthropic、OpenAI 这类公司在招投行的人、会计师、律师。你看到他们那么多垂直行业团队在发招聘广告,就是这个原因:你必须吃透这个领域的问题。如果你在某个主题上没有背景,也没真正下功夫补过课,那你基本不可能给它做出好的 eval(评测集)。好,假设我们已经理解了问题。我自认为对这个问题理解得还不错,这期播客听完你也会懂。那就进到 prompt 环节。还是那句话,我们要找的是那组「难度刚刚好」(Goldilocks)的 prompt。我习惯的第一步是先从简单的开始——你得先确认模型真的做得了这件事。另外,在 agentic(智能体式)场景里我说「模型」的时候,一般指的是模型加上外面那层 harness(脚手架 / 执行框架),这两个词我会混着用。这里说的是一个前沿模型,被包装成能调用这些工具、能做这些推理,
[18:43] Daniel
and it could come back with a solution. So for this problem of genome interpretation, I picked like one of the easiest genetic diseases possible and this is cystic fibrosis. This was something that we have known the genetic cause for quite some time and there are like canonical genes that cause cystic fibrosis. So to save you the effort of me googling for this, I just had the link right here and let's just go ahead and look at this. So what we see here is the canonical cystic fibrosis mutation in ClinVar which is an NIH database for uh a lot of genetic disease. And what we see here is it's got four stars and three stars. This really should be four and four. This is like the canonical uh genetic defect for cystic fibrosis. So I'm going to make sure that my agent can actually get this before I go forward. And so this is like the easy thing to start. So we're going to have our agentic genetics eval and we're going to say gene and we're going to say cftr2.
并且能给出答案。针对「基因组解读」这个问题,我挑了一种最简单的遗传病——囊性纤维化(cystic fibrosis)。它的遗传病因我们很早以前就搞清楚了,也有几个公认的致病基因。为了省得你们看我现场 Google,我提前把链接放这儿了,我们直接看。屏幕上这个就是囊性纤维化最经典的那个突变,在 ClinVar 里——ClinVar 是 NIH(美国国立卫生研究院)维护的遗传病数据库。你看它标着四星和三星,其实这里应该是四星四星才对,这就是囊性纤维化最标准的那个基因缺陷。所以在往下走之前,我要先确认我的 agent 真能把它找出来。这就是最容易的起点。我们来建这个 agentic 遗传学 eval,先填「基因」这一列,写上 CFTR。
[19:52] Daniel
This is the again canonical gene. And then we'll say variant. And what a variant is is how a particular gene is mutated. So right here this is again this is a somewhat niche eval but what this is saying is that on this particular transcript of CFTR at this position there is one base deleted and what that results in is the 508th fennel alanine deleted. So this is the actual variant that's going to exist in our eval. Sounds good. Okay. So what we see is that this is the exact variant that we care about. And what you'll notice is this is like very complex and nuanced. And this kind of comes back to the absolute first point I was making is you really need to know the space to do these evals. A lot of the eval are kind of like picked up like if you're trying to come up with an eval for Python coding like many of these are pretty good now. And for any of you who have used these models, saying Python coding is a solved problem is a little bit of a strong statement, but it is a very very well understood and well-characterized problem. So we will go to this and we're going to say okay so this is the thing we want to know and I should add a column here and this is the phenotype is cystic fibrosis. So what you see here is I'm starting to build a table of question and answer. So the question is I have this phenotype of cystic fibrosis which uh is is a lung disease and you know we'll describe exactly what happens there and then we have the genome of this patient and then we have the answer which is this variant. So let's walk through like how we would do that and when you're constructing these evals you're going to really really really use
这就是那个公认的致病基因。然后我们再加一列叫「变异」(variant)。所谓 variant,就是某个基因具体是怎么突变的。这里——再强调一次,这是个比较小众的 eval——它的意思是:在 CFTR 的这条特定转录本上,这个位置少了一个碱基,结果就是第 508 位的苯丙氨酸(phenylalanine)被删掉了。这就是我们 eval 里要用的那个真实变异。可以吧。好,这就是我们关心的那个变异。你会注意到,这东西非常复杂、非常讲究细节。这又回到我一开始说的第一点:要做这类 eval,你必须真懂这个领域。很多 eval 是被人随手拿来就用的——比如你想给 Python 编程做个 eval,现成的已经不少,而且质量都挺好。用过这些模型的人都知道,说「Python 编程已经被解决了」有点言过其实,但它确实是一个被理解得非常透彻、刻画得非常清楚的问题。我们回到表格,说:好,这就是我们想让它答出来的东西。这里我还得再加一列——表型(phenotype),填「囊性纤维化」。你看,我这是在搭一张「问题—答案」表。问题是:我有一个囊性纤维化的表型(这是一种肺部疾病,具体症状我们等会儿会描述),再加上这位病人的基因组;答案就是这个变异。我们来走一遍具体怎么做。而且你在构造这些 eval 的时候,会非常非常非常频繁地用到
[21:43] Daniel
genai a lot to construct them. So, what I'm going to do is now I'm going to go over to a terminal window I have open here. And I'm using codeex. Any of the tools will work. I have it on 53 spark low because I want it to be fast for this demonstration. But, uh, you know, you can use any any model, any anything you like. And for this task, this will be fine. So, what I have here, and I pre-baked some of these just to kind of make it go faster, but I want to walk through any step anyway, is I have actually two files here that are representing my genome. And what I want to do is create a synthetic version of this genome that has these variants that have the question and answer through this agentic flow that I want to get at. So what I'm going to do is I'm going to say please add and then this is going to be this variant to and this is a small variant. So it's going to come through here and create a new file in a new folder and we're going to call this uh dan cfive.vcf.gz GZ and we're going to have this be Dan CF live and we're actually asking AI to help us create the eval for AI. So we're going to do this and what the model is going to do is a variance file is literally just a text file. Actually we can see what it looks like here uh just for fun. So while while this is running, let's just take a look just so we know what these variant files look like. And uh this is this is fine. We don't need to show all of it. But what we basically see here is a chromosomes. So uh you might remember from things like 23 and me. We have 23 chromosomes. So this is one. This is the biggest one. And this is a position. So your chromosomes have
生成式 AI(genAI)来帮你造。所以现在我切到旁边开着的终端窗口。我用的是 Codex,其实换成任何工具都行。我把它设成了 5.3 Spark Low 这一档,因为演示需要快。当然你爱用什么模型都可以,对这个任务来说都够用。我这里有——有些东西我提前跑好了,为了让演示快一点,但每一步我还是会走一遍——这里其实有两个文件,代表的是我自己的基因组。我要做的是造一个合成版的基因组,把这些变异塞进去,这样就能通过这条 agentic 流程跑出我想要的那组问答。所以我输入:请把这个变异加进去——这是个小变异。它会在这儿新建一个文件夹和一个新文件,我们把它命名为 dan_cf5.vcf.gz。注意,我们其实是在让 AI 帮我们造给 AI 用的 eval。跑起来。模型要做的事其实很简单,因为变异文件(variant file)说白了就是个文本文件。趁它跑着,我们顺便看一眼它长什么样,就当解闷。让我们看看这些变异文件的样子。这样就够了,不用全看完。基本上,这一列是染色体——你可能从 23andMe 这类服务里记得,人有 23 对染色体。这个是 1 号,最大的那条。这一列是位置。因为染色体的
[23:49] Daniel
different number of bases. You might remember we have like three billion bases in our in our genome. And then these are swaps. So what we see is in this position a reference human has C and I have a CA here. So that means that um you know during some I inherited from my parents or maybe something that emerged during my development um I got a a base swapped here and then there's a bunch of metrics around quality how real it is. You you might remember you have two copies of each gene. So this is actually hetererozygous meaning only one copy is impacted. And you know, really, it's just a text file of all the letters in your alphabet. And what I'm saying is I want to add uh okay, this is still running. And if this is still running, in a while, I'll just use the pre-baked one. Is I just want to add this particular variant into my genome to see if our system can catch it. And this is what Codeex is doing right now. And I actually don't know why this is taking so long because this is like a oneliner.
碱基数量各不相同。你大概记得,人的基因组一共约 30 亿个碱基。再往后这些是替换。你看,在这个位置上,参考人类基因组是 C,而我这儿是 CA。也就是说,要么是我从父母那儿遗传来的,要么是我发育过程中出现的,反正这个位置的碱基被换掉了。后面还跟着一堆质量指标,说明这个结果有多可信。你可能还记得,每个基因我们都有两份拷贝。这一条是杂合的(heterozygous),意思是只有一份拷贝受影响。归根到底,它就是一个把你那套「字母」全列出来的文本文件。我让它做的就是加一条——好吧,它还在跑。要是再跑一会儿还不出来,我就直接用提前准备好的那份。我要做的只是把这个特定变异加进我的基因组,看看我们的系统能不能抓出来。这就是 Codex 现在在干的事。说实话我也不知道它为什么这么慢,这明明是一行就能搞定的活。
[24:53] Daniel
Um, but
呃,不过——
[24:54] Aakash
find the line I think or
我猜它是在找那一行吧,还是说……
[24:56] Daniel
Yeah. Yeah. Right. Right. I guess I'd probably I didn't want to make this demo too pre-baked. I thought about it. Should I just have a a Python program that just does all this for you? But then I'm like, that would not help the users at all because when they're constructing their own, they wouldn't know how to do it. So, just believe me that this will work. And for the sake of time, we'll go to the pre-baked ones. Mhm. So, uh, so basically what's going to happen is that codeex is going to add, this is in chromosome, where is it? I'm actually not sure. But in whatever chromosome this cystic fibrosis gene is in, it's just going to add one row and it's going to say we're going to have a deletion. So instead of having like a CA here, it'll just have a C. So like here's a deletion. You'll notice that we've we've lost an A. We went from TA to T. And then it's going to have a a fake cystic fibrosis patient. So now let's just check to see how this works. And we'll say uh use our agent. And again these are all agentic evals. And I'm using codeex as our agent but you could use anything clin open code your custom harness any way you could to get these agents to actually operate. And in fact even chat GPT and claude actually in the web UI they use agents right now.
对对,是这样。我本来就不想把这个演示做得太「预制」。我也纠结过:要不要直接写个 Python 程序,一键把这些都替你干了?但转念一想,那对观众一点帮助都没有,因为他们自己动手构造的时候还是不知道该怎么弄。所以你们就相信我,这一步是跑得通的。为了节省时间,我们直接用提前准备好的版本。基本上,Codex 要做的就是:在——这是几号染色体来着?我一时还真不确定。反正在囊性纤维化这个基因所在的那条染色体上,加一行,标明这里有个缺失(deletion)。原来这个位置是 CA,现在就只剩 C。这就是一个缺失,你能看到我们少了一个 A,从 TA 变成了 T。这样就造出了一个假的囊性纤维化病人。现在来看看效果。我们让 agent 上。再强调一遍,这些都是 agentic eval。我这里用 Codex 当 agent,但你用什么都行——Claude Code、OpenCode、你自己写的 harness,只要能让 agent 真正跑起来就行。事实上,现在连 ChatGPT 和 Claude 的网页版,背后用的也都是 agent。
[26:13] Daniel
They're not just model in, model out, they're model in, reasoning, tools, everything. Okay, cool. All right, so the agent finished here. So we see we added this in a record. Okay, and we see here it is. It's in chromosome 7. And you'll notice this is deletion. It says TCTT and instead it's a T. And so this means that this is like the canonical CF. So we're going to see if our agent is able to do it. And we're actually going to try a few different agents because one thing I mentioned is you want to score like you know 25 to 50% on these evals. You have to think about what tool are you using the you know mythos 5 for some biod defense thing then it's got to be really really hard or maybe this is something that for infrastructure reasons or cost reasons you need to use a very small model you need to use haik coup or something like that. So, we're actually going to try these simultaneously on a few agents and see what happens. So, we're going to say inside, what was this file we had? Uh, Dan CF live. Again, this is the one that we just made.
它们不是「输入进模型、输出出来」这么简单,而是输入、推理、工具调用,一整套。好,agent 跑完了。可以看到我们加进去了一条记录。就是这个——在 7 号染色体上。注意这是个缺失:原本是 TCTT,现在只剩一个 T。这说明它就是那个最经典的 CF 突变。接下来看看我们的 agent 能不能把它找出来。我们还会同时试几个不同的 agent。因为前面提过,你希望在这类 eval 上的得分落在 25% 到 50% 这个区间。你得想清楚自己用的是什么工具:如果是拿 Mythos 5 这种顶配去做生物防御相关的事,那题目就得出得非常非常难;也可能因为基础设施或成本的限制,你只能用很小的模型,比如 Haiku 这种。所以我们同时在好几个 agent 上跑,看看会怎么样。输入:在——刚才那个文件叫什么来着?dan_cf5。就是我们刚刚造出来的那个。
[27:22] Daniel
We have the genome of a patient suffering from, and let's just very quickly copy and paste some cystic fibrosis symptoms. Uh this one we suspect cystic fibrosis. Please find a genetic cause. And then we are going to we're actually going to copy this prompt so we can use it across multiple agents. So now we're going and we're we're we're checking GPT 3.5 codec sparklo. Uh we have a few other tabs open. So, I mentioned, you know, maybe we want to actually see if Haiku can do this. So, we'll just upload this. And again, we're going to the CF live. And let's also try Cat GPT 5.5 extra high. And I'm not going to use Pro because it will take too long. And this is also a relatively easy task. I suspect all the agents will get them. Okay, great. So we are cooking with haik coup and we are cooking with you can see what's all already happened with our first agent is after one minute of thinking you've actually find the cftr mutation pattern consistent with this deletion right this is the canonical cystic fibrosis gene so what this is telling us going back to our steps is this task is not too hard for these agents at least in the easy case
「这里是一位患者的基因组,他的症状是……」——我快速复制粘贴一段囊性纤维化的症状进来。「我们怀疑是囊性纤维化,请找出遗传学病因。」然后我把这个 prompt 复制一份,好让它能在多个 agent 上通用。现在跑的这个是 GPT-5.3 Codex Spark Low。旁边我还开着几个标签页。前面说了,我们也想看看 Haiku 行不行,所以把文件传上去,同样用那个 CF 文件。再试一个 ChatGPT 5.5 Extra High。我不打算用 Pro,太慢了,而且这个任务相对简单,我估计所有 agent 都能答出来。好,Haiku 那边开始跑了。你看第一个 agent 这边已经出结果了——思考一分钟后,它找到了 CFTR 上与这个缺失一致的突变模式。对,这就是囊性纤维化那个经典基因。回到我们的步骤,这说明什么?说明至少在简单档位上,这个任务对这些 agent 来说不算太难,
[28:55] Daniel
especially not with a powerful system like codeex. We'll see if haiku gets it. I suspect haiku will also get it. But you can see haiku even itself knows the cftr region is uh important. So while that's cooking let's go back to our next step. So again we start with something easy just to make sure it's possible. I knew that this was possible, but if you just told a layman and say, "Hey, could uh, you know, an AI agent find the canonical cause of cystic fibrosis inside a file with billions of variants?" They might say yes, they might say no, right? You just need to know. You need to kind of try it to get a sense of if it's possible. Quick thought experiment for you. Is there anything in this video you should be trying on your own? If there is, try it. Take a screenshot, post it on LinkedIn X, and tag me. I'd love to see what you're learning. Now, a quick word from our sponsors before we get into the back half of the pod. If you've worked at any company bigger than 30 people, you know this one. The CEO sets strategy. By the time it reaches the people actually doing the work, it goes through three or four layers of translation. Half of it gets lost and nobody finds out until the quarter is over. That's the problem AISO is built for. It's an AI operating partner for every manager and team.
尤其是用 Codex 这种强系统的时候。我们看看 Haiku 能不能拿下,我猜它也能。而且你看,连 Haiku 自己都知道 CFTR 这个区域很关键。趁它跑着,我们回到下一步。再说一次:先从简单的开始,只是为了确认这事儿做得到。我心里清楚它做得到,但你要是随便问一个外行:「一个 AI agent 能在一个装着几十亿条变异的文件里,找出囊性纤维化的经典病因吗?」他可能说能,也可能说不能。你必须自己确认,得亲手试一下,才能对「可不可行」心里有数。(以下为 Aakash 口播)给你留个小思考题:这期视频里有没有什么是你该自己动手试一试的?有的话就去试,截个图发到 LinkedIn 或 X 上,记得 @ 我,我很想看看你学到了什么。进入下半场之前,先插播一段赞助商内容。在超过 30 人的公司待过的人都懂这个场景:CEO 定下战略,等它传到真正干活的人手里,中间已经过了三四层转译,一半的信息在路上丢了,而且没人察觉,直到一个季度结束才后知后觉。AISO 就是冲着这个问题来的——它是给每一位管理者和每个团队配的 AI 运营搭档。
[30:08] Aakash
Connects to where work actually happens. the meetings, the messages, the docs, and it turns all that fragmented activity into a clear picture of execution. Managers get real coaching grounded in their team's actual work, not generic advice. Teams stay aligned with strategy as it changes, not as it was last quarter. And leaders see where execution is drifting in weeks, not in the post-mortem. One shared memory for the whole org. Everyone finally [music] working from the same picture. If you lead a team, check out ariso.ai/ashos. That's a riso. / a a kh I want to take a second to talk to you about the fourth cohort of LAN PM job. I trained 30 students in cohort 1, 50 students in cohort 2 and 75 students in cohort 3 and [music] I am bringing back the program for cohort 4. It starts in August and it lasts 3 months where you're going to have intense sessions a Monday morning session where I go over your resume, behavioral interviews, LinkedIn. On top of that, Bart Choworki is going to be teaching you the PM fundamentals in [music] 2026. how to write AI PRDS, how to AI prototype with cloud code, all of the key skills you need to freshen up your knowledge for this market. [music] And Ankut Romani is going to be teaching you AI product management. He is an AI product manager at Uber and he is going to teach you how to build AI features that actually work successfully. On top of that, Prasad Ready is going to be doing one-on- ones with you for mock reviews, LinkedIn review, candidate market fit review. So, it is a full package. It is three courses in one for one low fee. So join at landpob.com.
它接进工作真正发生的地方——会议、消息、文档,把这些碎片化的活动整合成一张清晰的执行全景图。管理者拿到的是基于团队真实工作的教练建议,而不是放之四海皆准的空话。团队对齐的是「此刻的战略」,而不是上个季度的战略。领导者能在几周之内就看到执行在哪儿偏了,而不是等到复盘会上才发现。整个组织共享一份记忆,所有人终于看的是同一张图。如果你带团队,去看看 aiso.ai/aakash。另外我想花点时间聊聊 Land PM Job 的第四期。第一期我带了 30 名学员,第二期 50 名,第三期 75 名,现在第四期回归了。八月开课,为期三个月,课程强度很高:每周一上午的直播课由我带你过简历、行为面试和 LinkedIn。除此之外,Bart Choworki 会教你 2026 年的 PM 基本功——怎么写 AI PRD、怎么用 Claude Code 做 AI 原型,全是这个就业市场里你需要更新的关键技能。Ankut Romani 会教 AI 产品管理,他是 Uber 的 AI 产品经理,会教你怎么做出真正跑得通的 AI 功能。另外,Prasad Reddy 会跟你做一对一:模拟面试复盘、LinkedIn 复盘、候选人与市场匹配度复盘。所以这是一整套打包方案,一份低价买到三门课。报名走 landpmjob.com。
[31:40] Daniel
Today's podcast is brought to you by Pendo, the leading software experience management platform. McKenzie found that 78% of companies are using Genai, but just as many have reported no bottom line improvements. So how do you know if your AI agents are actually working? Are they giving users the wrong answers, creating more work instead of less, improving retention, or hurting it? When your software data and AI data are disconnected, you can't answer these questions. But when you bring all your usage data together in one place, you can see what users do before, during, and after they use AI, showing you when agents work, how they help you grow, and when to prioritize on your roadmap. Pendo Agent Analytics is the only solution built to do this for product teams. Start measuring your AI's performance with agent analytics at pendo.io/acos. That's pendo.io aka. But then if you know that the easy thing works and you know we've already have early evidence the easy things works you have to like establish the ceiling is like what is is is the hard thing working cuz like if it's just totally saturated then like what's the point of even having eval this task is already solved. Um so I'm I'm going to show off something that's pretty hard to do today. And what we have here is a recent paper. So this is from uh last year and and I will make this bigger. The authors here are deciphering the diagenic architecture of congenital heart disease. So what does this mean? This means congenital heart disease is if you're a baby and you're born with problems with your heart and diagenic means it involves two genes. So single
本期播客由 Pendo 赞助播出——领先的软件体验管理平台。麦肯锡发现,78% 的公司都在用生成式 AI,但同样比例的公司反馈说:财务上根本没看到改善。那你怎么知道自己的 AI agent 到底有没有在干活?它是不是在给用户错误答案?是在制造更多工作而不是减少工作?它到底在提升留存,还是在伤害留存?当你的软件数据和 AI 数据是割裂的,这些问题你一个都答不上来。但如果把所有使用数据汇总到一处,你就能看到用户在用 AI 之前、之中、之后分别做了什么——agent 什么时候真的起作用、它怎么帮你增长、路线图上该优先做什么,一目了然。Pendo Agent Analytics 是唯一一个专为产品团队做这件事的方案。去 pendo.io/aakash 开始衡量你的 AI 表现吧,网址是 pendo.io/aakash。好,回到正题。当你已经确认简单任务做得通、已经拿到了早期证据,接下来你就得把「天花板」也立起来——也就是难的那头模型到底行不行。因为如果测试集已经完全饱和了(模型全做对),那这个 eval 还有什么意义?任务早就解决了。所以今天我要秀一个相当难的东西。我们打开的是一篇论文,去年发的,我把它放大一点。这篇的标题是《解析先天性心脏病的双基因(digenic)架构》。什么意思呢?先天性心脏病就是婴儿一出生心脏就有问题;digenic 的意思是它牵涉两个基因。相比之下,单
[33:20] Daniel
gene, single variant genetic diseases are actually sometimes a solved problem. Like with cystic fibrosis, not all cases, but many cases like this one are totally understand. Diagenic genetic diseases are like a very very new thing that people are studying. So this is like a very hard task to do and even though this is published and in theory an agent should be able to search the internet and find every publication and uh you know deeply understand all of this it's actually not that simple. They're not perfect and they actually need a lot of guidance which is why there are a lot of these companies including my own that are called like harness engineering companies or vertical AI companies or agentic AI companies because you need some specialized capability to be able to um have the LLM do stuff like this. So let's briefly return to our our agents and let's just make sure they got it. Okay. So, uh, codeex with GPT 5.3 says, okay, this, if you recall, this is the deletion that we added. Boom. I gave it the phenotype. I gave it the genome.
基因、单变异导致的遗传病,其实在某些情况下已经算是「解决了的问题」——比如囊性纤维化,不是所有病例,但像这类的很多病例机制已经完全搞清楚了。而双基因遗传病属于人们才刚刚开始研究的东西。所以这是个非常难的任务。虽然这篇论文已经发表了、理论上一个 agent 应该能上网搜到所有相关文献、深度理解全部内容,但现实没那么简单。它们并不完美,实际上需要大量的引导——这也正是为什么会有一堆公司(包括我自己这家)被叫作 harness engineering 公司(给模型搭脚手架的公司)、垂直 AI 公司或者 agentic AI 公司:你必须具备某种专门的能力,才能让 LLM 干成这种事。好,我们简单回到刚才那几个 agent,看看它们做出来没有。OK,Codex 配 GPT-5.3 的这个说……如果你还记得,这就是我们刚才手动加进去的那段缺失。搞定。我给了它表型,也给了它基因组。
[34:26] Daniel
This is correct. So, how would we mark this correct? We would actually probably have another LLM. I'm not going to do this right now for the sake of time. Just compare my scorecard. This is the correct answer with the response the model is giving right here. And then we can check Haiku. And even Haiku. Oh, wait. Did Haiku not get this? Okay. So, so this is interesting. This is actually harder than I would have thought. I would have expected Haiku to get this because this problem is so easy.
这个答案是对的。那我们怎么把它标成「正确」呢?实际做法多半是再上一个 LLM——为了节省时间我现在不演示了——让它拿我的评分卡(scorecard)里的标准答案,去和模型此刻给出的这段回复做比对。然后我们再看看 Haiku。连 Haiku 都……等等,Haiku 没做出来?有意思,这居然比我预想的还难。我本来以为 Haiku 肯定能做对,因为这道题实在太简单了。
[35:00] Daniel
But you'll notice what Haiku says is there's 48 variants spanning the gene. So, it's looking at the gene, but it fails to actually find the particular Oh, this is so interesting. It also hallucinates a hemisy large deletion. So coming back to this, coming back to our point is start with something easy. I thought I started with something easy here. It's a good thing I did this because if I were benchmarking highQ, this is too hard and I'd have to make it even easier. And the things I could do to make it even easier would be potentially uh, you know, limit the region of the genome of interest, give it more hints, maybe provide access to more external information more easily. But we can see that Haiku even fails this easy task. And I would be absolutely shocked if 5.5 did. Okay, it it it's not finished yet, but you can already see that it it found the correct answer. So if we were to score this, we could easily have an LLM say, okay, haiku, this is not correct. This does not match what I have in this table.
但你注意看 Haiku 说的是:这个基因上跨了 48 个变异。也就是说它确实盯到了那个基因,却没能定位到具体的那一个——哦,这太有意思了,它还幻觉出了一个「半合子大片段缺失」。所以回到我们刚才那句话:一定要从简单的开始。我以为我这里已经挑得够简单了。幸亏我先跑了这一步——因为如果我要给 Haiku 做基准测试,这道题对它来说就是太难了,我还得再往简单里改。能让它更简单的手段包括:把要关注的基因组区域限定得更窄一点、多给点提示、或者让它更容易拿到外部信息。总之我们看到,连这么简单的任务 Haiku 都做不出来。至于 5.5,它要是也做不出来我会非常震惊。好,它还没跑完,但你已经能看到它找到正确答案了。所以如果要打分,我们完全可以让一个 LLM 来判:Haiku 这个不对,跟我表格里的标准答案对不上。
[35:58] Daniel
This is correct and this is correct. So that's kind of and then in our in our spreadsheet we would just say you know haiku bad others good
这个对,这个也对。大概就是这样,然后在我们的表格里就记一笔:Haiku 差,另外两个好。
[36:11] Daniel
um and but now let's move on to something where we want to understand the hard cases and again I unexpectedly actually picked out a hard case for haik coup but this paper is quite challenging and I believe it is unlikely that any of the models will solve this. So in the supplementary information of this paper is a table and it is a list of patients and a proband is a medical term for the patient you're evaluating and it has diagenic causes for congenal heart disease. So this particular patient has ACACB I have no idea what this is some gene it's het meaning it only has one copy of this variant and myio CD which is also hat which is one carpy and these researchers discovered that the combination of these two diseases leads to congenital heart disease. So let's see if the models can figure this out. So what we'll do again is we'll go to our table and our phenotype.
接下来我们换个方向,去理解那些「难例」。刚才我其实是意外地给 Haiku 挑到了一道难题,但下面这篇论文是真的很有挑战性,我判断没有哪个模型能解出来。这篇论文的补充材料里有一张表,列的是一批患者——proband(先证者)是医学术语,指你正在评估的那个病人——表里给的是先天性心脏病的双基因致病组合。比如这个患者,一个是 ACACB,我完全不知道这是什么,反正是某个基因,标注是 het,意思是杂合、只带一份该变异;另一个是 MYOCD,同样是 het、也是一份。研究者发现,正是这两者的组合导致了先天性心脏病。我们来看看模型能不能推出来。还是老套路:回到我们那张表,先填表型。
[37:14] Aakash
Two diseases or is it two uh abnormalities in their DNA?
这是两种病,还是他 DNA 上的两处异常?
[37:18] Daniel
Yeah, that's a great question. It's one disease. It's a congenital heart disease and I don't know exactly which one it is from this paper. Um and you know we could read the paper and figure out exactly what the phenotype is, but it's some defect with the heart. And what's unusual about this and why this is hard is it's two hetererozygous variants on two different genes that is causing this single disease. So it's complicated. So our phenotype here is congenal heart disease and our gene here we have two of them. One is this guy and oops and the second one is this guy. And then the varants are these guys. And this is you'll notice the notation is a little bit different, but this is something that you'll just have to deal with in these evals is like, you know, no matter what you're doing because these tend to be in very technical specialized domains at this point. You know, no one wants eval for for boring stuff like ice cream flavors.
问得好。是一种病——先天性心脏病,具体是哪一种我从这篇论文里还没看出来。我们当然可以把论文读一遍、把表型确认到位,但反正是心脏上的某种缺陷。它不寻常、也正是它难的地方在于:是两个不同基因上的两个杂合变异共同导致了这一种病。所以很复杂。那我们这里的表型就写「先天性心脏病」,基因这一栏有两个:一个是这家伙,哎呀写错了,第二个是这家伙。然后变异是这几个。你会注意到这里的记法跟刚才那种不太一样——这也是做这类 eval 时你必须去适应的东西。不管你做的是什么,因为现阶段这些 eval 大多落在非常技术、非常专精的领域里。毕竟没人会去给「冰淇淋口味」这种无聊的东西做 eval。
[38:16] Daniel
You you just have to get comfortable with all this different mutations. And now let's go and let's try this again. So what I would do is I would say something like please add these to dan deep variant VCF. But I'm actually not going to do this because you already saw how this worked and basically how the VCF file was structured. It would add these two rows and for the sake of time I've already done it. But then let's go and let's check and see. Oh, how do we do on this use case? So in this case, I've already pre-baked it and I'm going to say this is Dan CHD and I'll say this contains the genome of a patient with congenal heart disease. Please identify the genetic cause. Okay. And while we're going to have this one running, we're going to try our other two agents just to see how they do. Um, we can almost guarantee that Haiku will not get this because it didn't get the much much easier task.
你只能让自己习惯这些五花八门的突变写法。好,我们再来一遍。我会写这么一句:请把这些加到 Dan 的 DeepVariant VCF 文件里。不过我这次不真的跑了,因为你刚才已经看过它是怎么工作的、VCF 文件大概是什么结构——它会往里加这两行。为了省时间我事先已经做好了。那我们直接来看:这个用例上模型表现如何?我这里是预先烤好的,我把它命名成 Dan_CHD,然后说:这是一位先天性心脏病患者的基因组,请找出遗传学病因。好。趁这个跑着,我们把另外两个 agent 也试一遍,看看它们怎么样。基本可以断定 Haiku 是做不出来的——那么简单得多的任务它都没做出来。
[39:37] Daniel
But for the sake of completion uh completeness, we will do this as well. And and then we'll do this with um GT5.5 extra high as well. And again, we would do
但为了完整起见,我们还是跑一遍。然后再用 GPT-5.5 的 extra high 档跑一遍。同样地,我们还会……
[39:52] Aakash
probably like some non-deterministic nature, right? Like do you need to like test the same model a couple times to just see if like maybe two out of three times it gets it right or is that not important?
……还有那个非确定性的问题对吧?就是你需不需要同一个模型多测几次,看看是不是三次里有两次能做对?还是说这个其实不重要?
[40:02] Daniel
Yeah, that's a really good point. Um so this is a question about sampling. Um, so sampling is actually really important. And you might remember um like all of this old research where you would basically sample for good traces. And this is kind of what like RL environments do is you do a roll out, you do a roll out, you do a roll out, and then you get the correct answer and boom, you give it a good reward for that. And the key thesis here is inside the weights of the model, the right answer might live there. it just might not get the right answer each time. So when you're doing these evals, you do want to try multiple times. In this particular case, I actually know from having done it that sampling has very little effect and it's essentially deterministic based on model capabilities. I've seen slightly the same model get to the same conclusion with slightly different approaches. But in general, sampling in my experience is less important than it used to be. where sampling used to be a big deal. Um like if you look at uh let's just look at this is a funny story. Um let's look at Gemini Ultra uh scorecard.
这一点问得非常好。这其实是采样(sampling)的问题。采样确实很重要。你可能还记得早年那一大批研究,基本思路就是不断采样、去捞出好的轨迹(trace)。这也差不多是 RL 环境在做的事:跑一次 rollout、再跑一次、再跑一次,直到拿到正确答案,然后 boom,给它一个正向奖励。这里的核心论点是:正确答案可能就存在模型权重里,只是它未必每次都能把它取出来。所以你做这些 eval 的时候,确实应该多试几次。不过就眼下这个案例,我是跑过的,我知道采样几乎没有影响,结果基本是确定性的,完全取决于模型能力本身。我见过同一个模型用略微不同的路径得到同样的结论。但总体上,以我的经验,采样的重要性比过去低了——过去采样可是件大事。比如说,讲个好玩的事,我们来看看 Gemini Ultra 当年那张成绩单。
[41:14] Daniel
So if you'll remember Gemini Ultra Oh wow. Did Google actually bury it? [laughter] Okay, here it is. This is So you'll remember way back in 2023 when people thought Google was kind of out of the AI race. I actually worked on this model so I know the story very well. Google released Gemini Ultra which was uh I believe it was a 660b dense model which was crazy back then. That was like one of the largest dense models ever trained and they released this scorecard and what you see is this. This was very controversial and this comes back to your question about sampling. If you remember MMLU, this used to be like the canonical benchmark for LLMs. And let's just go back to look at what a question is to remind you is just simple question answer. If you know the actual brightness of an object, its apparent brightness from location with no information, you can estimate this. Okay, models used to be bad at this, which is hilarious because this seems so distant right now. And Google wanted to be the best and GPT4 was the best at this point. Got 86.4% 4% on MLMU and it was on five shots meaning it had five samples and they picked the best one and that's why you see five shot three shot three shot 10 shot it really is like kind of like a way of cheating is like how many times can you sample um from this model and what you see is that Gemini Ultra actually had 32 shots so they got more shots on goal and actually now that I'm remembering this this actually might be pre-examples but if this is actually not the number examples in the context window and just the shots or or or the number of times sampled. It it it doesn't really matter for the sake
你要是还记得 Gemini Ultra……哦豁,Google 是不是把它给埋了?(笑)好,找到了,就是这个。回到 2023 年,那会儿大家都觉得 Google 在 AI 这场竞赛里基本出局了。这个模型我自己就参与过,所以来龙去脉我很清楚。Google 当时发布了 Gemini Ultra,我记得是个 6600 亿参数的稠密(dense)模型——在那个年代简直疯了,是当时训练过的最大稠密模型之一。他们发了一张成绩单,就是你看到的这张。这张表当年争议极大,而它正好接上你刚才问的采样(sampling)问题。你应该还记得 MMLU,它一度是评测 LLM 的标准基准。我们先回头看一道题长什么样,就是很简单的问答:如果你知道一个天体的实际亮度,再知道它在某个位置的视亮度,你就能估算出距离。就这类题,模型当年做得很烂——现在回头看简直好笑,感觉恍如隔世。当时 Google 特别想拿第一,而 GPT-4 是那会儿最强的,MMLU 拿了 86.4%,用的是 5-shot,也就是采样五次挑最好的那次。所以你会看到表里写着 5-shot、3-shot、3-shot、10-shot——说白了这有点像作弊:看你能从模型里采样多少次。而 Gemini Ultra 用的是 32-shot,等于射门机会多得多。哦,我现在回想起来,这里的 shot 也可能指的是 few-shot 示例数,而不是上下文里的示例数、单纯是采样次数。不过这对我要说的道理
[42:57] Daniel
of this argument, but the answers to MMLU might be inside the model weights, but it might just be not enriched enough in terms of the probabilities. So by sampling more times, you actually get a higher chance of getting the correct answer. So this used to be a really big thing back in the day. Today, I don't think this is a big thing. The labs don't really publish anymore. I don't think there's a lot known about this. In my personal experience, I've not found that running the same prompt through the model multiple times generates different answers. In fact, I don't know if I've ever seen that for this task, but it's a really really good thing to do and you should test for your use case. Little little side side uh conversation while we um look at the answers. And we've got all these cooking. These are all cooking. And we can see that 5.3 Spark finished first. This is actually why we did it. And what you'll notice is it's totally wrong. They found multiple variants. TBX1, my H cyst, JAG1. You'll notice these aren't even genes of interest for us. These variants aren't even relevant. No high confidence variants. Um, so what we've done here is we've done the second step in our process is we've established the floor.
没什么影响。真正的点在于:MMLU 的答案其实很可能就藏在模型权重里,只是在概率分布上没有被抬得足够高。所以你多采样几次,蒙对的概率自然就上去了。这事在当年是个大话题,今天我觉得已经不算什么了——各家实验室基本不再公布这些,外界对此了解也不多。就我个人经验,同一个 prompt 反复跑同一个模型,我没怎么见过答案会变;至少在这类任务上我印象里没遇到过。但这仍然是个非常值得做的检查,你应该针对自己的场景去测一测。……我们边等结果边闲聊两句。你看,这些都还在跑,都在锅里炖着呢。哦,5.3 Spark 最先跑完了——这也正是我们要试它的原因。你会注意到,它答得完全不对。它找出了一堆变异:TBX1、MYH6、JAG1。你会发现这些根本不是我们关心的基因,这些变异压根不相关,也没有高置信度的变异。所以到这一步,我们完成了流程里的第二步:把「地板」立起来了。
[44:10] Daniel
We found something easy, that cystic fibrosis gene. Now we found the ceiling is this model didn't get it. And plot twist, no model on the planet gets this without like a very very strong harness. Again, I'm working on that very strong harness. So, you know, we can make systems get this, but this is hard. And uh then you just kind of go through and it's almost like a binary search process where you say easy, medium, hard, and then just assemble a list of prompts. You know, you might have a hundred of these, which again would be phenotype gene, phenotype gene, phenotype gene, and then you understand how the models do on it. And then you're done. And then you have your eval. And then you understand what is good enough to actually ship product. So if you're scoring 50%, is that good enough to ship your product? Probably not. So then you look at which phenotypes am I better at? Maybe you put guard rails on the product to make sure that it only will answer the types of questions that it can get 80% on or something like that. That's a product decision. That's a product manager's decision to do that. And then you also hand all the hard ones to the research team and you say, "Hey guys, you didn't get this. Fix the model or fix the hardness and make sure that it can get these in the future so I can ship a product with these capabilities."
我们先找到了一个简单的——囊性纤维化那个基因;现在又找到了「天花板」:这个模型答不出来。剧透一下:这题地球上没有哪个模型能直接做对,除非配上一套非常非常强的 harness(外围脚手架系统)。我现在做的就是这套强 harness。所以我们是能让系统答对的,但这确实很难。接下来你就一路这么走下去,几乎就是个二分查找的过程:标出简单、中等、困难,然后攒出一份 prompt 清单。你可能会攒上一百条,每条都是「表型—基因、表型—基因、表型—基因」这样的配对,然后你就知道模型在上面表现如何。做到这儿就完事了,你的 eval 就有了。而且你还会知道:到底做到什么程度才够格把产品发出去。如果你只得 50 分,够发布吗?大概率不够。那你就再往下看:我在哪些表型上表现更好?也许你就给产品加护栏(guard rails),只让它回答那些能拿到 80 分的问题类型。这是个产品决策,是产品经理该做的决策。然后你把所有困难题打包丢给研究团队,跟他们说:「兄弟们,这些你们没做出来,去把模型修好,或者把 harness 修好,让它以后能答对,这样我才能把带这些能力的产品发出去。」
[45:24] Daniel
And uh not shockingly, Haiku is totally off base. Uh clearly HiQ is not good at all at this um task. Um so that's surprising. It missed the first one. Not surprising it missed this one. Um, and we again we have Oh my god. Okay, this is so interesting. Okay, so this is actually a good example of something that came out and sampled a second time and worked because I actually tried this. We were just talking about sampling. But we can see right now that GPT 5.5 extra high this time actually did identify this diagenic pairs and it did almost certainly find the paper. Yeah. So it did find the paper with all these diagenic pairs. So this is actually a very interesting reasoning trace where it was able to turn this congenital heart defect phenotype into a search for a very specific paper and then pull out the results from this paper. So actually this is quite impressive from uh GPT 5.5. But uh this is this is correct. So in this case, I would have to find an even harder one if I were benchmarking this model in particular. But this is basically the the key set of steps. And I don't think we have time to do a bunch more, but it's basically running through all the different types of scenarios and then coming up with prompts that will challenge the model but not totally stump the model. So yeah, with that, that's that's how you write an agentic eval. And here is two lines in our new one.
然后,毫不意外,Haiku 完全跑偏了。很明显 Haiku 在这个任务上一点也不行。它连第一题都做错了,这有点出乎意料;这一题做不出来倒是不意外。然后我们再看……我的天,太有意思了。好,这正好是个绝佳的例子——它在第二次采样时给出了不一样的结果,而且做对了,因为我之前试过一次。我们刚才不是还在聊采样嘛。你看,这次 GPT-5.5(extra high 推理档)真的识别出了这个双基因(digenic)组合,而且它几乎肯定是把那篇论文找出来了。对,它确实找到了那篇讲这些双基因组合的论文。所以这是一条非常有意思的推理轨迹:它把这个先天性心脏缺陷的表型,转化成了对一篇非常具体的论文的检索,再从论文里把结论提取出来。这一手 GPT-5.5 做得相当漂亮。总之这题它答对了。那么在这种情况下,如果我要专门给这个模型做基准测试,我就得再去找一道更难的题。但整套关键步骤基本就是这些。我们今天大概没时间再多演示了,但本质就是把各种类型的场景都过一遍,然后设计出既能挑战模型、又不至于把模型彻底难倒的 prompt。好,这就是写一个 agentic eval 的全过程。这里就是我们新建的这份 eval 里的两行。
[47:00] Aakash
Wow. So, it's just a spreadsheet. And the key thing here is the domain subject matter expertise. It's not like how it's written or anything like that. You're not giving us an EVEL template like we might have given people a PRD template before. It's really the subject matter expertise that's driving all of this.
哇。所以它就是一张表格而已。而这里真正关键的是领域上的专业知识(subject matter expertise),跟怎么写、格式如何都没关系。你今天没有给我们一个 eval 模板——不像我们以前会给大家一份 PRD 模板那样。真正驱动这一切的,是领域专业能力。
[47:19] Daniel
Yeah. Exactly. There are a lot of companies over the last, you know, n years who have tried to build better tools for evals. And I'm not saying that tools for evals don't need to exist. There's plenty of ways to improve, but when you're creating evals like this, it is literally just prompts, responses, and ways of scoring whether response is correct.
对,完全正确。过去这些年有一大堆公司都在做更好的 eval 工具。我不是说 eval 工具没有存在的必要,能改进的地方多得是。但当你像刚才那样去构建 eval 时,它字面意义上就只是:prompt、回答,以及一套判断回答对不对的打分方式。
[47:42] Aakash
Fascinating. So just to bring it all back full circle, if you're a PM who's never worked on an AI feature before, when would you be going through this eval process and when wouldn't you and how would you be using it?
太有意思了。那我们把话题收回来串一遍:如果你是个从没做过 AI 功能的 PM,什么时候该走这套 eval 流程,什么时候不该走,又该怎么用它?
[47:53] Daniel
Yeah, that's a a good question. So if you've never worked on an AI feature before, I would actually try to find somebody who has who can help you through this. This is deceptively simple. I made this really simple because we had 45 minutes today, but this and I don't know why it's so complex honestly. I had many many conversations with people about how to build evals but there's just something kind of like taste based or or nuanced about how to build them and you know it is what it is but fi find somebody who can help you but I would say you start from the beginning is like you are building an AI feature like I don't know I'm just making it up you work at Pinterest and you want to have uh better image generation that doesn't look like slop that actually pleases users so it's like an image generation feature you need to think about from day zero. What does success look like? What is unique about those Pinterest users? What do they want to see? And you need to translate that. You can't just write down a PRD and say they want beautiful kitchens. You have to explicitly define what a beautiful kitchen is and not in words in examples and a way to score those examples. And I am not in the image generation space, so I don't exactly know what that looks like. But there are many many many examples like this where you need to understand what the user wants and translate translate that into prompts and responses and way to score those responses. And that is the first thing you should do when you're starting to build a new AI feature.
嗯,好问题。如果你从没做过 AI 功能,我建议你先找一个做过的人带你走一遍。这事看着简单,其实有陷阱。我今天把它讲得特别简单,是因为我们只有 45 分钟。说实话我也不知道它为什么会这么复杂——我跟很多人聊过怎么做 eval,但它就是有一种偏「品味」、偏微妙的东西在里面,事实就是这样。所以,去找个能帮你的人。不过我会说,你要从最开头就开始想。比如你在做一个 AI 功能——我随便编一个——假设你在 Pinterest 工作,你想做更好的图像生成,生成的东西不能是那种一眼 AI 垃圾(slop),而是真能取悦用户的图。那么从第零天你就得想清楚:成功长什么样?Pinterest 的用户有什么独特之处?他们想看到什么?然后你得把这些翻译出来。你不能只在 PRD 里写一句「用户想要漂亮的厨房」就完事,你必须明确定义什么叫「漂亮的厨房」——而且不是用文字定义,是用示例(examples)定义,再配一套给这些示例打分的方法。我不是做图像生成的,所以我不太清楚具体该长什么样。但这类例子非常非常多:你得搞明白用户到底想要什么,再把它翻译成 prompt、回答,以及给回答打分的方式。这就是你着手做一个新 AI 功能时,第一件该干的事。
[49:20] Aakash
Okay. So just like you have ramped up your expertise in the genomic space if you were tackling that problem you'd go learn, you'd go talk to people who have built image and eval oh this is how I build an LLM judge that generates images of this type. And then that would really be the basis for your eval.
明白。所以就像你为了做基因组这块把自己的专业度硬拉起来一样——如果你去啃那个问题,你会去学,会去找那些做过图像 eval 的人聊:哦,原来做这类图像的 LLM-as-judge(用大模型当裁判打分)是这么搭的。然后这些东西才真正成为你 eval 的基础。
[49:39] Daniel
Yep, that's correct.
对,没错。
[49:39] Aakash
Okay. Wow, there's so much so many layers. I've done like five or six episodes on eval, but I think this was one of the most tactical that really helped me understand how things change and I think that's a function of your experience, which I wanted to talk about for a little bit. So, your eval piece crossed my radar. I think another really interesting piece you wrote about was product management at Meta versus Google. You've worked on Gemini, you've worked on Llama, you've seen both of these cultures. What really is the difference between product management at Meta versus Google?
好。哇,这里面层次真的太多了。关于 eval 我大概做过五六期节目,但我觉得这期是最落地的一期,真正让我明白了事情是怎么变化的。我想这跟你的经历有关,这块我也想聊一会儿。你写的那篇 eval 的文章进了我的雷达,另一篇我觉得特别有意思的,是你写 Meta 和 Google 的产品管理有什么不同。你做过 Gemini,也做过 Llama,两边的文化你都见过。Meta 和 Google 的产品管理,真正的差别到底在哪?
[50:15] Daniel
Yeah. So, I would caveat that and say I wrote this like two and a half years ago and I was at Google three and a half years ago, I think. So, a lot has changed. When I was at Google, Google was a dead company. I think the stock fell to like $80 and I think it's you know 300 or 400 right now and uh they've really changed how they think about things. I'm a boomerang at that so I think I've spent seven years there in in in total and you know I saw everything from Cambridge Analytica lows to highs of like Llama 3 really wowing people to lows of Llama 4 disappointing. So I I saw a a very large spectrum and what I would say like my key takeaways for what uh Google versus Meta was like is Meta is like a much much more aggressive culture uh in many ways. I think that it comes from like the founder leadership of Mark Zuckerberg is he is the last man standing. Well, I guess besides Elon, but he is the last man standing who's got like, you know, the founder leading a fan company who has utter and absolute control who's going to do what he wants.
嗯。我得先加个前提:那篇文章我是两年半前写的,而我离开 Google 大概是三年半前的事了,所以变化很大。我在 Google 那会儿,Google 简直是家「死掉的公司」,股价跌到了 80 块左右,现在大概是三四百块,他们的思路确实彻底变了。我在 Meta 算是「回锅」的(boomerang,离职后又回去),前后加起来待了七年,什么都见过了——从剑桥分析(Cambridge Analytica)那种低谷,到 Llama 3 惊艳众人的高峰,再到 Llama 4 让人失望的低谷。所以我经历的跨度非常大。要说 Google 和 Meta 最大的差别,我的核心体感是:Meta 的文化在很多方面都要激进得多、攻击性强得多。我觉得这跟 Mark Zuckerberg 的创始人式领导有关——他是最后一个还在位的人。好吧,除了 Elon 之外,他是最后一个还在亲自掌舵 FAANG 级别公司的创始人,手握绝对控制权,想干什么就干什么。
[51:20] Daniel
And sometimes it's really empowering because he says, "This is super important to me. You have all the resources in the world and you should go do it." And sometimes it's like not what you want. For example, I worked on Llama before. Llama had problems. I think the main problem was actually how it was evaluated. Huh, funny. Those emails are important. If you want to read about that story, Google it. I had nothing to do with that and I I loved working on Llama before and Mark just said, "You guys all suck. You need to go find new jobs." So basically the whole Llama team is gone because of you know Mark's his decisions. So I think that you know that that cuts both ways. I'd say my overall preference is for like a very like high conviction founder company. Um I actually have like incredible respect for Mark. I've only you know met him a couple times and uh every time has been just like wow this is like a really smart guy but it creates a lot of problems too right because you know Google is much more consensus driven I think Google has much weaker product management function at least it did when I was there so uh you know Google is considered to be more of an engineeringled company whereas meta is more of a productled company at least that historically has been the case and um it was a really great experience working at both places I think I've learned a lot for both places. I left both places with a lot of friends and uh yeah, if you want to read like kind of this blog post actually went pretty viral. If you want to read like kind of a interesting snapshot of what it was like in say 2023 between both of the
有时候这特别提气,因为他会说:「这件事对我极其重要,公司所有资源都给你,放手去做。」但有时候,结果就不是你想要的了。比如我以前做过 Llama。Llama 出了问题,而我认为最核心的问题恰恰出在它是怎么被评估的。挺讽刺的吧。那几封邮件很关键。想了解那段故事的话,自己去 Google 一下。那事跟我没关系,我以前在 Llama 做得很开心,然后 Mark 一句话:「你们这帮人全都不行,去找新工作吧。」于是整个 Llama 团队基本上就散了,就因为 Mark 的那些决定。所以我说这事是双刃剑。总体上我还是更偏爱那种创始人信念极强、说了算的公司。我其实对 Mark 有非常高的敬意——我只见过他几次,但每次都是「哇,这人是真聪明」的感觉。可这种模式也会带来一堆问题。Google 就更讲共识驱动,我觉得 Google 的产品管理职能弱得多,至少我在的时候是这样。所以外界一般说 Google 是工程驱动的公司,Meta 更偏产品驱动,至少历史上一直是这个样子。在这两家的经历我都觉得非常棒,都学到很多,离开的时候也都留下了一大堆朋友。如果你想看,我写过一篇博客,传播得还挺广的——想看 2023 年前后这两家公司各自是什么状态的一个切片,
[52:47] Daniel
places, uh you know, give it a read.
可以去读一读那篇。
[52:48] Aakash
Highly recommend it to everybody. As you guys can see, I'm itching to ask many more questions. So Daniel, we're going to need to have you back. Before you go, tell us a little bit about your startup.
强烈推荐大家都去看看。大家也看得出来,我还有一肚子问题想问,所以 Daniel,下次我们得再请你来一期。在结束之前,跟我们讲讲你的创业公司吧。
[52:57] Daniel
Oh, cool. Yeah, so I've started a company called Gamoff Labs. And as I hinted at during these evaluations, the core problem I want to solve is to make it much much easier to get whole genome sequencing into every single NICU in the entire world. This is the absolute gold standard of helping sick babies. Um there's overwhelming clinical and economic evidence that it's effective, but the problem is it's just too damn hard and expensive. So where you see this used is in places like Stanford, Boston Children's, uh CHOP and these are the absolute top facilities in the world and I want to see them in, you know, rural Arkansas, rural India, you know, rural China, all the places where um this this kind of life-saving technology is not being harnessed. And my key thesis is that a lot of the human work involved in interpreting these genomes could be augmented by AI. And we've already shown that using the system that we've built, we can identify variants that have never been discovered before.
好啊。我创办了一家公司叫 Gamoff Labs。正如我在讲 eval(评估)的时候暗示过的,我最想解决的核心问题是:让全基因组测序(whole genome sequencing)能真正低门槛地进到全世界每一家 NICU(新生儿重症监护室)。这是救治重症新生儿的绝对黄金标准。临床证据和经济性证据都压倒性地证明它有效,问题在于它实在太难、太贵。所以现在你只能在 Stanford、Boston Children's、CHOP(费城儿童医院)这些地方看到它在用,全是世界最顶尖的机构。而我希望它出现在阿肯色州的乡下、印度的乡下、中国的乡下,出现在所有这类救命技术还没被用起来的地方。我的核心判断是:解读这些基因组过程中大量的人工工作,是可以被 AI 增强的。用我们自己搭的这套系统,我们已经证明能识别出此前从未被发现过的变异。
[54:01] Daniel
We've actually allowed one family to have a child and they they couldn't before because they didn't know. Again, this is small scale. We started this five weeks ago. So, you know, but one person is is really crazy to have that impact on their life. And um yeah, like it's a deeply deeply missiondriven thing. I think it's very very interesting technically because it's all about building the best agentic harnesses. It's all about understanding how AI can help with biology. And if you're interested in joining me on this journey, we are hiring right now. Um we're very small team. Uh ra raised our preede round and are basically planning on building like the operating system for rare disease and genomic medicine. And I couldn't be more excited to wake up to work on this every morning. And I would love uh I would love it if you would reach out if you're interested.
我们其实已经让一个家庭敢生孩子了——他们以前不敢,因为不知道病因。当然,规模还很小,我们五周前才开始。但哪怕只改变了一个人的人生,那种影响也已经挺不可思议了。这件事的使命感非常非常强。技术上我也觉得极其有意思,因为它本质上就是在打造最好的 agentic harness(智能体的工程骨架),就是在搞清楚 AI 到底怎么帮上生物学的忙。如果你有兴趣加入这段旅程,我们现在正在招人。团队非常小,刚融完 pre-seed 轮,接下来基本上是要造一套罕见病和基因组医学的「操作系统」。每天早上醒来能干这件事,我兴奋得不能再兴奋了。如果你感兴趣,非常欢迎来找我聊。
[54:47] Aakash
Wow. So a lot of you guys I know at least in my audience you want to become that AIBM at meta Google this is often the next step after that. So if you know we always say the grass is greener at some point this is where I've seen those AIPMs at Meta and Google go just like Daniel into starting their own companies and that's actually the cool thing is it helps prepare you for that. You can see how his own evals and deep AI knowledge has now applied to his startup. Daniel, thank you so so much for lending your expertise today.
哇。我知道我的听众里有不少人想成为 Meta、Google 那样的 AI PM,而这往往就是之后的下一步。我们老说「别人家的草更绿」,但走到某个节点,我看到的那些 Meta、Google 的 AI PM 都跟 Daniel 一样——出去开自己的公司。有意思的地方在于,前面那段经历恰恰是在为这一步做准备。你能看到,他做 eval 的功夫和对 AI 的深度理解,现在全都用到了他自己的创业上。Daniel,非常非常感谢你今天来分享你的专业经验。
[55:19] Daniel
Yeah, and I want to leave with just one parting thought is the metas and the Googles and all the other large companies have to reinvent themselves right now in the age of AI. If you're inside these companies, it is very very interesting to see how this classic consumer software building factory has changed. But if you come and you do a startup or you start your own thing, you get to build the future from scratch. And sometimes that's actually easier.
好的。我最后再留一个想法:Meta、Google 以及其它所有大公司,在 AI 时代都必须重新发明自己。如果你在这些公司里面,看着这套经典的「消费级软件生产流水线」是怎么被改写的,会非常有意思。但如果你出来创业、自己做,你就能从零开始造未来。而有时候,那反而更容易。
[55:46] Aakash
We'll leave it there. See you all in the next episode. I hope you learned as much from today's episode as I did. If you can do one thing that's totally free that would help the show, it would be to check that you're following on Apple and Spotify podcasts. Check that you've left ratings and reviews on those platforms. Check that you're subscribed on YouTube. Leave a like and a comment on this video. And then share it with your friends. [music] We're trying to make better and better podcasts. After 2 years, we think we've gotten something pretty good going. So, let us know what we can do to make it even better, who else we should interview, and we will put on the best shows [music] we possibly can. Finally, don't forget my offer for the bundle. You get an entire year of my paid newsletter, plus my favorite AI tools, Bolt, new, Air Table, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arise, and Mobin.
今天就聊到这儿,我们下期再见。希望你今天的收获和我一样多。如果你愿意做一件完全免费、又对这档节目最有帮助的事,那就是去 Apple Podcasts 和 Spotify 上确认你已经关注了,顺手留个评分和评论;在 YouTube 上确认你已经订阅,给这期视频点个赞、留条评论,然后分享给你的朋友。[音乐] 我们一直想把播客做得越来越好。做了两年,我们觉得已经摸出点门道了。所以也告诉我们还能怎么改进、接下来该采访谁,我们会尽全力做出最好的节目。[音乐] 最后,别忘了我那个打包优惠:你能拿到我付费 newsletter 的一整年订阅,外加我最喜欢的一批 AI 工具——Bolt.new、Airtable、Speechify、Descript、Magic Patterns、Linear、Dovetail、Arize 和 Mobbin。
[56:36] Aakash
That's $27,000 worth of value for just $150. So check that out at bundle.ashg.com if it interests you. And I can't wait to share our next episode soon.
加起来价值 27,000 美元的东西,只要 150 美元。感兴趣的话去 bundle.aakashg.com 看看。我已经迫不及待想跟大家分享下一期节目了。