The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
频道: AI Engineer
视频: https://www.youtube.com/watch?v=ObTPqBGsEbA
原文语言: en
统计: 共 22 轮 · Sandipan 19
[0:07]
[music]
[音乐]
[0:15] Sandipan
All right. Um, thank you for joining my session.
好的。嗯,感谢大家来听我这一场。
[0:17]
[applause]
[掌声]
[0:20] Sandipan
Thank you, man. Uh, I'm Sandy. Uh, I'm a technical lead uh, for data and AI at Databricks. Um, prior to working in Databricks, I worked in Amazon Web Services uh, for 5 years as a principal architect for data and AI. Uh, in the past few years, I worked extensively uh, building and scaling data and AI platforms using distributed systems and technology. And in the past couple of years, specifically, I've been working with customers trying to figure out what we do with this new AI technology. Uh, when I say new AI, AI has been here for a long time, but we all started experimenting quite exponentially uh, in in the past couple of years, right? And I have learned a great deal of lessons on from building demos and how to take those demos to production working with different customers in uh, B2B software industries and then uh, regulating industries like financial uh, services. So, in this session, I want to share a playbook, a framework that I put together from lessons that I have learned working in the trenches uh, that you can take and apply on when you think about how to put your AI systems into production. And I think this session is nicely placed in the afternoon because what you can do now is in this framework, you can fit the different um, knowledge, the knowledge that you've gathered attending these different sessions throughout the day and see where they fit in each of these you know, elements in the framework. So, when I started 2 years ago, this is the pattern I noticed in every customer conversation, right? So, everyone wanted to do something with AI. Uh, there was immense pressure from the top to do something, to build a demo. And every conversation started with let's choose the model, right? And it was nobody's fault because the market was like that. We were talking about models, the models were new technology for us, right? And every conversation started, shall we use GPT? Shall we use Claude? You know, there was huge debate with with within organizations. Then, you would choose a model, you'll build some features, offsets of over what features to build for that application.
谢谢,谢谢大家。我叫 Sandy,是 Databricks 数据与 AI 方向的技术负责人。在加入 Databricks 之前,我在 Amazon Web Services 待了 5 年,做数据与 AI 的首席架构师。过去这些年,我做了大量用分布式系统和相关技术去构建、扩展数据和 AI 平台的工作。而最近这两年,我主要是在和客户一起摸索:这套全新的 AI 技术到底该怎么用。当然我说「新 AI」,AI 其实早就存在了,但真正开始爆发式地大规模尝试,也就是这两年的事,对吧?在和 B2B 软件行业以及金融服务这类强监管行业的各种客户打交道的过程中,我从「怎么做 demo」到「怎么把 demo 推上生产」学到了非常多的教训。所以这一场,我想分享一套 playbook——一个我在一线摸爬滚打总结出来的框架,你可以直接拿去用,去思考怎么把你的 AI 系统真正推到生产环境。我觉得这一场安排在下午挺合适,因为你现在可以做的事,就是把你今天听各个分场所积累的各种知识,往这个框架里一一对号入座,看它们分别落在框架里哪个环节。两年前我刚开始做这件事时,我注意到几乎每一次客户对话都是同一个套路:所有人都想拿 AI 干点什么。上面给的压力巨大,必须做点东西出来,必须做个 demo。而且每次对话的开场都是「咱们先选个模型吧」,对吧?这也不能怪谁,因为当时整个市场就是这样。我们满嘴都是模型,模型对我们来说就是新技术。每次开口都是:用 GPT 吗?还是用 Claude?组织内部为这个吵得不可开交。然后你选定一个模型,做几个功能,又为了「这个应用到底该做哪些功能」反复扯皮。
[2:30] Sandipan
Uh, you would build that in a controlled environment, so predictable data sets, you know, um, limited scenarios, and then it looked great as a demo, and then leadership would get happy, they would sign it off, and they'll put it into an environment in a production environment. Then, after a few weeks, people would start asking questions that what the hell is AI doing? Right? Why is it not answering the questions the way we expected it to answer when we were doing the demos? Uh, it would result in not only less you know, no realization in return on investment, but also loss of money and effort in building these demos that can never scale to production. Throughout these uh, meetings, I gathered three insights that connect to everything that you we are talking about when thinking about taking uh, AI to production. The first one is the observability gap, right? When we use AI and put it into production, if we can't see what it is actually doing, if we can't trace every decision that it's making, it's no use in production. Second is the evaluation gap gap. A lot of these conversations that we were doing, we were not actually thinking about what is what is that one thing that we are measuring. Yes, we talk about accuracy, we talk about latency, we talk about groundedness, but we were not defining what is that exact thing like that matters to the business, and how can we build a system that can continuously measure that, whether it's improving, whether it's not improving, like what what is that system that we need to build. And that was that evaluation gap that I noticed. And the third is the governance gap. Like we were not actually thinking what happens when AI fails in production. Who's accountable? Who do I go to when something happens at 3:00 a.m. in the morning, right? Who needs to own the data assets that feed some AI responses? What happens if AI you know, uh, talks um, uh, um, nonsense to a customer, right? What what happens, right? So, there is no accountability, no governance around it.
你会在一个受控环境里把它搭出来——数据集是可预测的,场景是有限的,于是 demo 看起来棒极了,领导一高兴就拍板放行,把它扔进生产环境。结果几周之后,大家开始发问:这 AI 到底在干嘛?为什么它回答问题的方式,跟我们做 demo 时期待的完全不一样?这种情况不仅意味着投资回报根本无从谈起,更意味着你在这些永远没法上生产的 demo 上白白搭进去的钱和精力都打了水漂。在这一场场会议里,我总结出了三个洞察,它们贯穿了「把 AI 推上生产」这件事的方方面面。第一个是可观测性鸿沟(observability gap)。当我们把 AI 投入生产,如果看不到它实际在做什么,如果无法追踪它做的每一个决策,那它在生产里就毫无用处。第二个是评估鸿沟(evaluation gap)。我们当时那么多对话,其实从来没认真想过:我们到底要衡量的是哪一个核心指标?是的,我们会聊准确率、聊延迟、聊 groundedness(答案有没有事实依据),但我们从没把那个「真正对业务重要的东西」定义清楚,也没想过怎么搭一套系统去持续衡量它——它到底在变好还是没变好,我们究竟需要搭一套什么样的系统。这就是我注意到的评估鸿沟。第三个是治理鸿沟(governance gap)。我们其实根本没去想:AI 在生产里出事了怎么办?谁来担责?凌晨三点出了岔子,我该去找谁?喂给 AI 答案的那些数据资产,该由谁来负责?如果 AI 对客户胡说八道了,会怎么样?到底会怎么样?所以根本没有问责机制,也没有任何治理。
[4:39] Sandipan
And these three insights led me to build a framework on how I think AI should be taken to production, and this has been implemented across multiple customer organizations, and I think this is something that you can pick up from here. These are the five pillars, and these are absolutely what you need to think about even before starting a project, right? Then you start build them gradually, preferably in sequence, but in real life, I know that this sequence don't work, but these are the pillars that you have to know about and you have to think about when start building. First one is evaluation. Before touching any code, before discussing about any models, any features, you have to think about when we build this system, how do we measure? What does success look like, and what is that system that will help us continuously measure what success looks like for us? Second is how do we trace each and every decision that AI makes. It's not only important for the performance of the AI system, it is also important for the regulators. In Europe or in a lot of companies, especially in regulated industry, you cannot even onboard AI into production without having tracing and observability in place. So, this is a must-have. The third is the data foundation, right? Uh, I I I I think of data foundation in in two ways. One is the question data, so that is basically the data needed for the AI to answer questions that users ask to it. So, it could be your pre-training data, post post-training data, data that you use APIs to hook onto and and get to the answer that the user needs. The other one is the tracking data, related to the tracing data in observability, but when you think from the data foundation and and uh, data strategy perspective, this needs to be handled in this pillar because you need a whole data strategy now with tracing data, especially when you run hundreds of agents in your organization. Fourth is orchestration. One agent would work pretty well. You don't need to think about orchestration. But when you onboard five agents, the the complexity increases exponentially, right? You will have multiple coordination patterns between these agents, they will need to talk to each other in multiple different ways, they will need to each wait for each other's responses, there's a lot of complexity that comes in. And that's where orchestration patterns and thinking about how you will orchestrate your agents in a particular system becomes really important.
正是这三个洞察,促使我搭出了一套框架——关于我认为该如何把 AI 推上生产。这套框架已经在多家客户组织里落地,我觉得是你今天可以直接带走的东西。这就是五大支柱,而且这些是你在项目还没启动之前就必须想清楚的,对吧?想清楚之后再逐步去建,最好是按顺序来。当然现实里我知道这个顺序根本走不通,但不管怎样,这几根支柱是你开工时必须知道、必须想到的。第一根是评估。在动任何一行代码之前、在讨论任何模型和功能之前,你就得想:我们建这套系统,到底拿什么来衡量?成功长什么样?又该有一套什么样的系统,帮我们持续衡量「对我们而言成功是什么」?第二根是怎么追踪 AI 做出的每一个决策。这不仅关乎 AI 系统的性能,也关乎监管方。在欧洲,或者在很多公司——尤其是强监管行业——你要是没把 tracing 和可观测性做到位,AI 根本就不可能上生产。所以这是个硬性必备项。第三根是数据底座(data foundation)。我看数据底座一般分两块。一块是「问题数据」,也就是 AI 回答用户提问所需要的数据,可能是你的预训练数据、后训练数据,也可能是你通过 API 挂上去、用来拿到用户所需答案的数据。另一块是「追踪数据」,它跟可观测性里的 tracing 数据相关,但从数据底座和数据战略的角度看,它得在这根支柱里专门处理——因为现在你需要一整套针对 tracing 数据的数据战略,尤其当你在组织里跑着成百上千个 agent 的时候。第四根是编排(orchestration)。一个 agent 的话运转得挺好,你压根不用想编排的事。可一旦你上了五个 agent,复杂度就指数级飙升,对吧?这些 agent 之间会有多种协同模式,它们得用各种不同方式互相通信,还得彼此等待对方的响应,一大堆复杂性就涌进来了。这时候,编排模式、以及「在一套特定系统里你打算怎么编排你的 agent」就变得至关重要。
[7:06] Sandipan
Fifth is governance. This is where you think about what happens when something fails. Who's accountable? How do we govern data? How do we secure it? How do we secure our systems? How what do we make sure how do we make sure that no one injects into our agent and leads uh, to you know, misbehavior, right? Or loss of reputation. So, in the rest of the session, I will dive a bit deeper into each of these pillars and tell you how how you can think about when you start working with them, right? The first one is evaluation. Evaluation is basically specification for your AI system. You define success. As I mentioned, it's not like talking about accuracy. You have to define it with numbers, like what accuracy is is is good for your business use case, right? Define it in numbers. Uh, define what kind of you know, false positives you can handle. What should be the deflection? So, this is this is an example from a a retail chatbot, right? A banking chatbot where when you implement a chatbot with an AI agent, one of the main goals is to deflect simple queries um, to the agent so that a human agent don't need to uh, deal with them, right? And so, you need to uh, you need to track those queries and track those numbers and put that system in place. Second is building those test test cases, like the evaluation data set. You've heard about golden data sets in evaluation. Talk with the domain experts and find what is actually happening in real life on the ground. Like what answer would a support human support agent um, give to a customer on a particular question. Collect those information. What happens in gray areas, in edge cases, like what happens when a human sees a customer asking a confusing question, right? Collect those into a data set, and then automate your AI testing, right? So, you put a question to AI, it answers, take that answer, compare against the test set, and automate this whole pipeline so that when you put AI in production, that pipeline can actually take live responses and evaluate against the test data set that you're building, and then give you the result in terms of how AI is performing against those numbers and the goals that you've defined.
第五根是治理(governance)。这一块你要想的是:出事了怎么办?谁来担责?我们怎么治理数据?怎么保护数据?怎么保护我们的系统?怎么确保没人往我们的 agent 里注入恶意指令、害它乱来,进而砸了我们的口碑?所以接下来这一场剩下的时间,我会把这五根支柱挨个往深里讲一点,告诉你上手的时候该怎么思考。第一根,评估。评估本质上就是你 AI 系统的规格说明书(specification)。你要定义成功。就像我说的,不是嘴上聊聊准确率就完了,你得拿数字把它定下来——比如对你的业务用例来说,准确率到多少才算好,对吧?用数字定义它。再比如你能容忍什么样的误报(false positive)、分流率(deflection)应该是多少。这里有个例子,来自一个零售聊天机器人——其实是个银行聊天机器人。当你用 AI agent 做一个聊天机器人时,一个核心目标就是把简单的咨询分流给 agent 去处理,这样人工坐席就不用去管这些了,对吧?所以你得追踪这些咨询、追踪这些数字,并把这套系统建起来。第二件事是构建那些测试用例,也就是评估数据集。你应该听说过评估里的「黄金数据集」(golden dataset)。去跟领域专家聊,搞清楚现实中真实情况是什么样的——比如人工客服面对某个具体问题时会怎么回答客户。把这些信息收集起来。还有灰色地带、边缘情况下会发生什么——比如人工坐席碰到客户问了个含糊不清的问题时会怎么处理?把这些都收进一个数据集,然后把你的 AI 测试自动化。你给 AI 抛一个问题,它作答,你把这个答案拿来跟测试集比对,并把整条流水线自动化,这样当你把 AI 推上生产,这条流水线就能真的接住线上的实时响应、拿它跟你正在构建的测试数据集去比对,再把结果反馈给你——告诉你 AI 相对你定下的那些数字和目标,表现到底如何。
[9:17] Sandipan
When we talk about evaluation, there are three main layers that I see appear across organization, and this is an architectural decision that you need to make when you build these evaluation systems, right? The first layer is deterministic. These are the easy stuff, like you know, checking formats, you know, checking email formats, phone formats, the regular expression things that we have already been doing with our coding systems, right? The uh, the the other other is like you know, you could use a classic ML models for name entity recognition to for intent classification, for understanding what is first name, last name, PII detection, etc. So, the these these are easy stuff, cheap stuff, you should get them out of the way. We have already been doing this for years. The second layer is the non-deterministic semantic stuff, all right? This is where groundedness comes in. This is where we implement technologies like LLMs judges. We all know what LLMs judges are, right? Everyone? Okay, I see a lot of nods. So, um Again, this This is a pretty simple version of how a uh how a prompt would look for an LLM as a judge. Um So, with LLM as a judge, you you you use a separate LLM from the primary LLM to judge the response of the primary model. And when you do that, you tell the secondary, the judge model, on how it should uh judge the primary model's output. So, it could be around safety, groundedness, you know, relevance to the answer, etc. etc., right? Again, that can feed from a lot of uh the evaluation data set that you have created, right? To look at what are the expected answers, and then it can check against that. This is a sample prompt on how these things work, but I'm sure you've attended some of these sessions where you've seen vendors doing this automatically at scale. Uh for example, in Databricks we in MLflow you'll find automatic LLM as judge, uh where you can create these custom LLM as judges that run automatically on traces. That's your second layer. The third layer is behavioral, right? This is where uh you think about a tool calls, like is our agents calling the right tool? Are they getting into loops? So, for example, um you know, the first layer you you can have a user ask a question, "What is my account balance?"
说到评估,我在各个组织里反复看到三个主要的层次,这也是你在搭评估系统时必须做的一个架构决策,对吧?第一层是确定性的(deterministic)。这些都是简单活儿——比如校验格式,校验邮箱格式、电话格式,就是我们写代码时早就在做的那些正则表达式那一套,对吧?再比如你可以用经典的 ML 模型做命名实体识别(NER)、做意图分类,去识别什么是名、什么是姓,做 PII(个人隐私信息)检测等等。这些都是简单、便宜的活儿,应该先把它们清理掉。这些我们已经做了好多年了。第二层是非确定性的、语义层面的东西,对吧?groundedness(答案是否有事实依据)就出现在这一层。我们也是在这一层引入「LLM as a judge(用 LLM 当裁判)」这类技术。大家都知道什么是 LLM as a judge 吧?都知道吗?好,我看到好多人点头。嗯,这里是一个相当简化的版本,展示给 LLM 裁判用的 prompt 大概长什么样。所谓 LLM as a judge,就是你用一个跟主模型分开的、独立的 LLM,去给主模型的回答打分。你这么做的时候,要告诉那个次级的裁判模型:它该按什么标准去评判主模型的输出。可以是围绕安全性、groundedness、答案的相关性等等,对吧?而它同样可以从你之前构建的评估数据集里取料——看看期望答案是什么,再拿来比对。这里是一个示例 prompt,展示它大概怎么运作。不过我相信你也听过一些分场,看到厂商已经能把这事自动化、规模化地做了。比如在 Databricks 的 MLflow 里就有自动的 LLM as a judge,你可以创建这些自定义的 LLM 裁判,让它们在 trace 上自动跑起来。这就是你的第二层。第三层是行为层(behavioral),对吧?这一层你要想的是工具调用(tool call)——我们的 agent 调对工具了吗?它们会不会陷进死循环?举个例子,在第一层,你可以让用户问一个问题:「我的账户余额是多少?」
[11:31] Sandipan
And you could go and check that, okay, this there is no deterministic problem with it. The seman- the agent answered right, that, "Okay, your account balance is this many dollars." And that was right, and you can see this is this is right, but when you go into the behavioral checks, you will see that the agent uh actually made three calls to the database to find that answer. Right? And that is because it was doing making duplicate calls for whatever reason. Calls failed, you know, calls did not work, it went and retried and stuff like that. Now, three API calls in demo environment is fine, but in production, when you get thousands of queries from users every day, and there's like duplication in API calls, that's an expensive operation. And that's where you need to think about behavioral evaluation. And this layer is very, very important. I see a lot of organizations, a lot of teams miss them when when talking about this. The second layer is observability, right? Uh and in this pillar, what we're talking about tracing, right? So, you collect [snorts] all the decisions that an agent is making. So, I want to explain this with a scenario here, right? And this is a scenario from an actual project I worked on with a banking uh retail retail banking chatbot. Now, obviously, if you've seen tracing data, it's not as beautiful as this slide, right? So, I've simplified it and made it beautiful for this slide. But what this slide says is basically, a user comes in and says, uh "You know, I have been charged an overdraft fee, can you waive it for me?" Because the user thinks that the customer thinks that that is not legitimate. So, the agent does an intent classification, and you all you know about this because you've enabled observability, you're capturing traces, and you're actually seeing what the agent is doing, right? What AI is doing. Intent classification, it is done, this it took this many seconds, this was confidence score. Then it goes and connects to the customer's account, maybe in a database, a customer database, call calls an API, connects to the customer database, gets the account details.
然后你去查,发现 OK,这里没有确定性层面的问题。语义上 agent 也答对了,说「好的,您的账户余额是多少多少美元」。这是对的,你一看也确实对。可一旦你进到行为层的检查,你就会发现:这个 agent 为了拿到这个答案,居然往数据库发了三次调用。对吧?原因是它不知为何在重复发请求——可能是某次调用失败了、没成功,它就跑去重试,诸如此类。现在,在 demo 环境里发三次 API 调用没什么大不了,但到了生产环境,每天有成千上万条用户查询涌进来,再加上 API 调用里的这种重复,那就是一笔昂贵的开销了。这正是你必须考虑行为层评估的地方。这一层非常非常重要。我看到很多组织、很多团队在谈这件事时都把它漏掉了。第二根支柱是可观测性(observability),对吧?这根支柱里我们谈的是 tracing。你要把 agent 做出的所有决策都收集起来。我想用一个场景来讲清楚这点,对吧?这个场景来自我做过的一个真实项目——一个零售银行的聊天机器人。当然,要是你见过真实的 tracing 数据,就知道它绝对没有这张幻灯片这么漂亮,对吧?我为了这张幻灯片把它简化、美化了一下。但这张图想说的本质是:一个用户进来说,「嗯,我被收了一笔透支费,你能帮我免掉吗?」因为这个客户觉得这笔费用收得不合理。于是 agent 先做一次意图分类——这些你全都看得到,因为你启用了可观测性、在捕获 trace,你能真切看到 agent 在做什么、AI 在做什么。意图分类,完成了,花了这么多秒,置信度是这么高。接着它去连接客户的账户,可能是在某个客户数据库里,调一个 API,连上客户数据库,拿到账户明细。
[13:25] Sandipan
It retrieves policy documents. It checks from a rag vector database, um what is uh what is the policy around overdraft, right? Is what the customer claiming is legitimate? So, it checks for policy documents. Then it goes and does a reasoning on what should be uh you know, responded to the customer, and then it does some final guardrail checks, and responds to the customer. Now, if you did not set up a system that helps you look visualize all of these traces, when the customer comes to you and raises a dispute, you have no way to check what the AI did. Right? You have nowhere to go, and you end up saying that I don't have have any idea. Let's Let's give the customer a discount or something, and then make them happy. So, this is why you need this, and this is why regulators are are are basically mandating, because otherwise there's no production system if you cannot do this kind of stuff. So, this is where um you know, you you you detect this the example that I gave around duplicate API calls. This is where you start detecting this stuff. So, when you when you enable these traces, you can actually go and see duplicate calls, and then take relevant actions based on that. Not only that, you can actually do that in online monitoring. So, when it's happening in production, in on you can set up online monitoring, and at that point, if it is doing duplicate calls, you can apply fallback strategies. Or even if it is doing a call that is failing, you can actually go and apply a strategy where it will say, "Okay, go and retry for three times, not more than three times. If it if it is more than three times, then report somewhere, or pass it to a human to take some action." The third pillar is the most important pillar, in my opinion, is the data data data foundation, right? Uh in my typical project projects, I spend 60% of my time, uh and I see I see a lot of organizations spending a lot of time here, because no one expected agents to come suddenly in the market and start querying data.
然后它去检索政策文档。它从一个 RAG 的向量数据库里查,关于透支的政策到底是怎么规定的,对吧?客户主张的这个诉求合不合理?于是它去查政策文档。接着它做一番推理,想清楚到底该怎么回复客户,然后再跑一些最终的护栏(guardrail)检查,最后回复客户。现在,如果你没搭一套能帮你把所有这些 trace 可视化出来的系统,那么当客户找上门来发起争议时,你根本无从去查 AI 当时到底做了什么,对吧?你压根没地方可查,最后只能说我也不知道——要不给客户打个折什么的,把人哄高兴算了。这正是你为什么需要它,也正是为什么监管方基本上把它列为强制要求——因为你要是做不到这种事,就根本谈不上生产系统。这也是你开始检测出问题的地方,比如我刚举的那个重复 API 调用的例子。当你启用这些 trace,你就真的能去看到那些重复调用,然后据此采取相应措施。不仅如此,你还能在线上监控里做这件事。当它正在生产环境里跑的时候,你可以搭一套在线监控,一旦它在做重复调用,你就能套用兜底(fallback)策略。或者即便它发起的某个调用一直失败,你也可以套一个策略,让它「好,去重试,但最多三次,别超过三次;要是超过三次,就上报到某个地方,或者交给人工去处理」。第三根支柱,在我看来是最重要的一根,就是数据底座(data foundation),对吧?在我那些典型项目里,我有 60% 的时间都花在这上面,而且我看到很多组织也在这儿投入了大量时间——因为谁也没料到 agent 会突然冒出来、开始直接查数据。
[15:27] Sandipan
Data was always built for humans, and humans are always forgiving. You find the wrong data in a report, you just go and ask someone to correct it. Agents don't forgive you, right? Agents will go, find it wrong, they'll give you the wrong answer confidently. Right? And you wouldn't know what's happening. And this is why data quality, setting the right data strategy, has become so important for enterprises now. I divide it into two sections. One is the question data, as I was explaining, like data needed for actually serving the AI's uh outcome. And the other one is the tracking data. This is the observability data, the tracing data I was talking about earlier. You need a proper plan on how you collect this tracing data, and how you serve it to auditors, to regulators, to do online monitoring, to run LLM as judges on the tracing, and everything else, right? So, there it needs a proper strategy on how you structure the schema and everything on the tracing data. Um On Databricks, um we create a robust data foundation for our customers using uh some of the technologies that we provide. If you don't know Databricks, Databricks has been built on some open-source technologies like Apache Spark, MLflow, and Delta Lake. Uh we provide a bunch of capabilities on top of it. So, the blue layer at the bottom is basically your cloud storage. Databricks works on the three major clouds, Google, AWS, Azure. Okay. I thought it was for me. So, so once you once you store raw data on your cloud storage, uh the data is then um um we we we bring in a a layer called the Delta Lake layer, which uh which basically brings in database-like properties on top of your raw data. So, you have got images, text files, video files, or whatever. We we help you create this um you know, uh table-like structure on top of it using manifest files, right? And and we help you to uh incrementally load data, do all of those um data management tasks in a structured way. On top of that, we bring in Unity Catalog, which is a data catalog. Uh with Unity Catalog, you can centrally apply permissions on top of the data.
数据一直以来都是为人构建的,而人总是宽容的。你在报告里发现一处数据错了,顶多去找个人改一下就完事。可 agent 不会这么宽容你,对吧?agent 会拿着那条错数据,理直气壮地给你一个错误答案,而你根本不知道出了什么岔子。这就是为什么数据质量、为什么把数据战略制定对,对今天的企业来说变得如此重要。我把它分成两块。一块是「问题数据」,就像我前面解释的,是真正用来支撑 AI 产出结果的数据。另一块是「追踪数据」,也就是我前面讲的那些可观测性数据、tracing 数据。你需要一套像样的方案,规划好怎么收集这些 tracing 数据,怎么把它们交付给审计方、交付给监管方,怎么用它做在线监控,怎么在 tracing 上跑 LLM 裁判,等等等等。所以这里需要一套像样的策略,去设计 tracing 数据的 schema 之类的一切。在 Databricks 上,我们用我们提供的一些技术,为客户搭建一个稳健的数据底座。如果你不了解 Databricks——它是构建在一些开源技术之上的,比如 Apache Spark、MLflow 和 Delta Lake。我们在这之上提供了一大堆能力。最底下那一层蓝色的,基本上就是你的云存储。Databricks 在三大主流云上都能跑:Google、AWS、Azure。好。我还以为(铃声)是在叫我呢。所以,一旦你把原始数据存到云存储上,我们接着会引入一个叫 Delta Lake 的层,它本质上是在你的原始数据之上叠加类似数据库的特性。所以哪怕你有的是图片、文本文件、视频文件之类的,我们都能帮你用 manifest 文件在它之上构建出这种类似表(table)的结构,对吧?我们还能帮你增量地加载数据,用一种结构化的方式去完成所有那些数据管理任务。在这之上,我们再引入 Unity Catalog,它是一个数据目录(data catalog)。有了 Unity Catalog,你就能在数据之上集中地施加权限。
[17:43] Sandipan
You can um you can uh share the data using uh Delta Sharing, but also uh what happens with Unity Catalog is uh you you can enable discovery and um you know, um ownership, metadata tagging capabilities at the catalog level. What that means is, when you apply table a description, column description, uh tag columns uh like PII columns with metadata, it becomes really easy for AI to then get that context when it queries these tables on top of Unity Catalog. So, everything is governed at one layer through Unity Catalog, and on on top of that we bring in different uh applications. So, whether it's AI through Mosaic AI, so to build LLM, tune LLM, or even build AI applications, we bring in uh data warehousing capabilities, BI capabilities, and um uh and some of the other text-to-SQL capabilities. We have got Genie that uh helps you write natural language to do SQL querying, etc. And one application of that in the observability and tracking tracking data, as I was showing, is is this. So, basically, think about when I was talking about the tracking data strategy. Organizations, especially enterprises, will not be running AI in just one framework. They'll be using different frameworks, CrewAI, LangChain, etc. etc. They'll be using different cloud platforms. And once they do that, you need a centralized layer of collecting that tracing data, so that you can serve sev- several use cases on the right hand side. So, whether it's for operational dashboarding, for first line support, uh a lot of these uh first line um first line of defense teams need health monitoring uh sort of dashboards, right? These teams can also write SQL using Databricks Genie to do text-to-SQL. But they can also build Databricks apps using coding agents uh to create common workspaces or custom uh UIs that customers might need for different uh different use cases. And then we've got Agent Bricks and MLflow that serves you uh LLM out of the box LLM as judges, and uh proactively monitor a The idea is, no no no matter where your AI runs, you can create this kind of strategy bringing in data in one common place and serving uh different teams from one shared location.
你可以用 Delta Sharing 来共享数据,但 Unity Catalog 的另一个好处是,它能在 catalog 这一层开启发现、归属、元数据打标这些能力。这意味着,当你给一张表加上描述、给字段加上描述、给比如 PII 字段打上元数据标签之后,AI 在查询 Unity Catalog 上的这些表时,就能非常轻松地拿到这些上下文。所以一切都在 Unity Catalog 这一层统一治理,在它之上我们再接入各种应用。无论是通过 Mosaic AI 来做 AI、构建和调优 LLM,还是搭建 AI 应用,我们都接入了数据仓库能力、BI 能力,还有一些 text-to-SQL 的能力。我们有 Genie,能帮你用自然语言来做 SQL 查询等等。而它在可观测性和追踪数据上的一个应用,就像我刚才展示的,就是这个。回到我之前讲的追踪数据策略:企业组织绝不会只在一个框架里跑 AI,他们会用 CrewAI、LangChain 等各种框架,会用不同的云平台。一旦这样,你就需要一个集中的层来收集这些追踪数据,这样才能服务右侧那几类用途。无论是给运维做仪表盘、给一线支持用,很多这种一线防线团队都需要健康监控类的仪表盘对吧?这些团队也可以用 Databricks Genie 写 SQL,做 text-to-SQL。他们还能用编程 agent 来搭建 Databricks 应用,为不同场景创建共享工作区或客户需要的定制 UI。然后我们有 Agent Bricks 和 MLflow,给你开箱即用的 LLM as judge,主动做监控。核心思路是:不管你的 AI 跑在哪里,你都能用这套策略,把数据汇聚到一个统一的地方,再从这个共享位置服务不同的团队。
[20:05] Sandipan
The fourth pillar is multi-agent orchestration patterns. As I said, one agent is good, multiple agents increases complexity. That's where you start thinking about, okay, what pattern is good for my use case. The first one I describe here is the orchestrator worker pattern. Where you have one orchestrator which orchestrates all the work, which controls all the work from a centralized plane, and then distributes this work to different agents based on their specialized skills. And then every request goes through the orchestrator, so you have got central control. If something goes wrong, you can go to the orchestrator logs and look into them and see what has happened. Right? So, that's the orchestration data uh pattern. There is this choreography pattern where each agent is independent, they're autonomous, they don't depend on an orchestrator. All of them talk to a message bus and they listen to the events that they are interested in. Right? So, think about agents that are independent of each other, right? They can run parallelly. So, they are not sequential, like one agent is not dependent on another. So, they run parallelly, they listen to the message bus for the for the events that they are interested in. Maybe it's a trigger for, let's say, a mortgage application, and it says, uh you know, uh the mortgage application agent uh one of the agent uh looks customer details, right? The other agent looks at approval details and everything else, right? They can work in parallel, and the advantage it brings you is the latency is reduced because they are not dependent on an orchestrator and sending messages back and forth. Right? So, this is the choreography pattern. And the third one is human in the loop, which is where when an agent crosses a threshold or serves below threshold a confidence threshold, then a human is called in the workflow to look into the pattern uh so, look into the looking into what the agent has done and then take action based on that. I have done a deep dive video on multi-agent orchestration pattern uh for the online track of this conference. Uh it's already on YouTube, so you can look into it. I talk about the real implications of when you think about multi-agent patterns.
第四根支柱是多 agent 编排模式。就像我说的,一个 agent 不错,多个 agent 就会增加复杂度。这时候你就得开始想:哪种模式适合我的场景。我这里讲的第一种是 orchestrator-worker(编排者-工作者)模式。你有一个 orchestrator,由它在一个集中的平面上统一编排、统一掌控所有工作,再根据各个 agent 的专长把任务分发下去。每个请求都要经过 orchestrator,所以你有了中心化的控制。一旦出问题,你可以去翻 orchestrator 的日志,查清楚到底发生了什么对吧?这就是编排模式。还有一种叫 choreography(编舞)模式,每个 agent 都是独立、自治的,不依赖 orchestrator。它们都连到一条消息总线,只监听自己关心的事件。所以你可以把它们想成彼此独立的 agent,它们能并行跑,不是串行的,一个 agent 不依赖另一个。它们并行运行,在消息总线上监听自己关心的事件。比如触发了一笔房贷申请,房贷申请的某个 agent 去看客户信息对吧?另一个 agent 去看审批信息等等。它们可以并行工作,这样带来的好处是延迟降低了,因为它们不依赖 orchestrator 来回传消息对吧?这就是 choreography 模式。第三种是 human in the loop(人在回路),就是当某个 agent 超过某个阈值、或者置信度低于某个阈值时,工作流里就会拉一个人进来,看看 agent 都做了什么,再据此采取行动。我为这次大会的线上专场做过一期关于多 agent 编排模式的深度讲解视频,已经发在 YouTube 上了,你们可以去看。我在里面讲了当你认真考虑多 agent 模式时,真正会遇到的那些现实问题。
[22:11] Sandipan
One is uh state management, the other is fault tolerance, like what happens when things fail, like how do you manage them? I talk about different patterns. And then talk about how how you think about scaling them in large scale on enterprises. Pillar five is governance, right? Now, here I'm not talking about data governance at all. That's given, we need that, right? From AI perspective, what what what are we thinking about? Regulatory, right? Audit trails, have we got the trail of every action, every user connection, every request, everything that happens in the system? Are we capturing everything? Are we doing pre-validation of personal information? Are we using name entity recognition? The the easy stuff, the rejects and all of those things, right? In our example, the work that I was doing with the customer that I mentioned, we already detected 47 PII breaches during the testing phase by applying this layer. So, that's that's really important. Fourth is um prompt versioning. You have to treat prompt versioning as change management in enterprise grade solution. It cannot be just change to a prompt and commit to get. It has to be is it has to go through proper change management processes as you do with code. So, basically treating prompt as code. Third is model change management. So, as models change, the model providers upgrade these models, you have to have a system to understand whether that upgraded model will be good for your use case, for your data. Right? Model providers you will put evaluation benchmarks on three uh benchmark uh boards, but those are not really useful when you put them in your context, in your enterprise. So, that's where these evaluation data sets come in handy, where you try these different models on this evaluation data set and try to understand which one performs better. And that management needs to be done because from a risk perspective, you cannot really rely on one single model. You have to have the flexibility to switch to different models and also test them on your own data.
一个是状态管理,另一个是容错——就是出故障的时候怎么办、怎么处理。我讲了不同的模式,还讲了在大型企业里怎么把它们做大规模扩展。第五根支柱是治理。注意,我这里完全不是在说数据治理,那是默认就该有的,我们当然需要对吧?从 AI 的角度,我们要想什么?合规对吧?审计轨迹——我们有没有记录下每一个动作、每一次用户连接、每一个请求、系统里发生的每一件事?我们是不是都捕获下来了?我们有没有对个人信息做前置校验?有没有用命名实体识别?那些基础的东西、那些该拒绝该过滤的东西对吧?在我们的例子里,就是我提到的跟那个客户做的项目,光是加上这一层,我们在测试阶段就已经检测出 47 起 PII 泄露。所以这真的很重要。第四是 prompt 版本管理。在企业级方案里,你必须把 prompt 版本管理当成变更管理来对待。它不能只是改一下 prompt 然后 commit 到 Git,它必须像对待代码那样走完整的变更管理流程。本质上就是把 prompt 当代码来管。第三是模型变更管理。随着模型更新、模型厂商升级这些模型,你得有一套机制去判断升级后的模型对你的场景、对你的数据是不是更好对吧?模型厂商会在那三大榜单上放评测基准,但把它们放到你自己的语境、你自己的企业里时,那些基准其实没多大用。这时候这些评测数据集就派上用场了——你拿不同的模型在这个评测数据集上跑一跑,看哪个表现更好。这种管理是必须做的,因为从风险角度看,你不能真的只押在单一一个模型上。你得有灵活切换到不同模型的余地,还得能在自己的数据上测试它们。
[24:07] Sandipan
That management needs to be done. Uh in Databricks, uh we have taken all of these these points, these pillars that I've been talking about into Agent Bricks. We are building Agent Bricks to make uh all of these operations out of the box for you, uh so that it's easy to implement production grade AI applications on uh on in in your enterprises. So, I wanted to quickly touch upon a case study, just to give you a uh a flavor of how these things go, right? So, when I was working with this client um they were a retail banking they were building a retail banking chatbot you know, one and a half 18 months ago. Uh their their problem the the problem they wanted to solve is they had got around 20,000 odd calls per month from customers on their chatbot. They wanted to deflect they they they saw that there were like 60% of them were simple queries, what is my account balance, you know, what do I do with my overdraft and all of those stuff, like that can be answered simply. So, they wanted to the reliance on human agents for those answers. So, they identified those queries and they wanted to automate them. Right? They spent around 85K in 6 months doing a POC which did not succeed. When we got involved, we found those insights that I was talking like no one knew why things were failing when it was in production when when when when it was in production. No one could actually measure why why it's not succeeding and no one could actually understand who is accountable for what when things go wrong. Right? So, the goal we set for them is AI agent handles 60% of user queries, right? Which were simple user queries and then a way to identify and track them. The key difference in this project that we when we did is that we selected the model in week seven, like in a eight weeks POC. Right? And this is how it turned out. For the week one and two, we built the evaluation layer. We collected 200 cases on their actual human agents answering to their customers on simple queries and understand how they are responding to them.
这种管理是必须做的。在 Databricks,我们把我一直在讲的这些点、这些支柱,全都做进了 Agent Bricks。我们打造 Agent Bricks,就是为了把所有这些操作做成开箱即用,让你在企业里实现生产级 AI 应用变得简单。我想快速讲一个案例研究,让大家感受一下这些事情实际是怎么走的。我跟这个客户合作的时候,大概一年半、18 个月前,他们在做一个零售银行的聊天机器人。他们想解决的问题是:他们的聊天机器人每月大概收到两万多次客户来电(咨询)。他们想做分流,他们发现其中差不多 60% 都是简单问题,比如我的账户余额是多少、我的透支该怎么处理之类的,这些都能简单回答。所以他们想减少这些回答对人工坐席的依赖。于是他们识别出这些问题,想把它们自动化对吧?他们花了大概 8.5 万美元、做了 6 个月的 POC,结果没成功。等我们介入时,就发现了我刚说的那些洞察:当系统上了生产,没人知道东西为什么会出错;上了生产之后,没人能真正衡量出它为什么不成功,出了问题也没人能搞清楚到底谁该为什么负责对吧?所以我们给他们定的目标是:让 AI agent 处理掉 60% 的用户问询——也就是那些简单问询,并且要有办法把它们识别和追踪出来。这个项目里我们做的关键不同点在于:我们是在第七周才选模型的,整个 POC 是八周。最后效果是这样的——第一周和第二周,我们先搭好评测层。我们采集了 200 个真实案例,看他们自己的人工坐席是怎么回答客户那些简单问询的,搞清楚他们是怎么应对的。
[26:21] Sandipan
We created that database. Then we defined the success metrics. What does success look like to you? So, out of let's say 100 queries, you need 60 queries or the 60% of the queries that are simple queries to be uh to be handled by the agent, right? They needed some sort of accuracy. So, 85% it was around 85% accuracy target. They needed latency, all of the operational targets that you need. They were there. Then we created this automated evaluation pipeline for them. And what what I mean by that is an automated system where you can capture a user's uh a user's question and the AI agent's response. You take that, compare that against your evaluation data set. You rate that, and if the rating is below certain threshold, you get it checked by a human. And you if if something goes wrong, you make sure that you find the solution. So, it could be a change to the prompt, it could be change to a tool calling system or something else. Once you have done that, you add that test case in the test data set. So, that when it happens next time, the test cases cases catch them. So, the the the the summary of that story is that your evaluation data set is a living system. You start with 200, maybe there is no correct number here, but once you start, as you start building in production, this is a living system. This will keep growing. And the and the bigger it grows, the better your system will be. In the second week, we talked we thought about the foundational layer, right? So, the the question data, we thought we thought that, okay, if you have to call the database, have you got the API connections right? Have you got a system in place that can trace the API connections? Are those secure, right? We were not talking about MCP at that time. Right? It was just direct API calls to database to run queries. Have you got the distributed storage? Have you got the Have you Are you collecting traces? And this is where when when we started testing after building these systems, we could catch those duplicate API calls, right? We could catch why customer satisfaction was dropping and stuff like that.
我们建好了那个数据库,然后定义成功指标。对你来说,成功是什么样子?比如 100 个问询里,你需要其中 60 个、也就是那 60% 的简单问询交由 agent 来处理对吧?他们需要某种程度的准确率,是 85%,准确率目标大概在 85% 左右。他们需要延迟达标,所有你需要的运维指标都得有,这些都到位了。然后我们给他们搭了这套自动化评测流水线。我说的意思是:一套自动化系统,能捕获用户的问题和 AI agent 的回答,拿过来跟你的评测数据集做比对,给它打分,如果分数低于某个阈值,就交给人来核查;一旦发现哪里出了问题,你就要确保找到解决办法。可能是改 prompt,可能是改工具调用机制,或者别的。改完之后,你把那条测试用例加进测试数据集里。这样下次再发生时,这些测试用例就能把它逮住。所以这个故事的总结是:你的评测数据集是一个活的系统。你从 200 条起步,这里也许没有什么标准数字,但一旦你开始在生产中迭代,这就是个活系统,它会不断长大;而且它长得越大,你的系统就越好。第二周,我们考虑基础层。比如那些问题数据,我们想:好,如果你要调数据库,你的 API 连接对不对?有没有一套机制能追踪这些 API 连接?它们安全吗对吧?那会儿我们还没在谈 MCP 对吧?就只是直接调数据库的 API 来跑查询。你有没有分布式存储?有没有在采集追踪数据?正是在搭完这些系统、开始测试的时候,我们才逮住了那些重复的 API 调用对吧?才搞清楚客户满意度为什么在下滑之类的问题。
[28:24] Sandipan
And then comes in week seven to eight, we started talking about models. Now that we had the evaluation data set, we could run different models on that data set to see the responses, compare them against the expected responses, and calculate a number on on on the accuracy, right? That helped us to decide which model to use. Now, that decision didn't take long, right? We in as I explained in the introduction, like we spent weeks debating on which model to use, but when you took the other approach, you can actually uh do that in a very quick way. So, once once that's done, we we stitched everything that I was talking about around observability, evaluation, the layers of evaluation. Once we had that system that can make AI visible, measurable, and accountable, that's when we started launching it to production. And that's when uh so, this is this is the result uh six weeks post launch, uh we we calculated the operational metrics, of course, you know, the accuracy, the deflection rate, the response time, uh the customer CSAT. But what's important here is in few weeks time, when uh there was a problem with uh so, one of the one of the things that happened was that the bank changed some uh interest rate related policies. So, when that when they changed the policy, they actually sent emails to customers or notifications in the application in the in the in the mobile banking app about the policy change. But when the customers came and queried on the chatbot for further questions, they couldn't get the right answers and they were like putting thumbs down on the answers. So they were getting this feedback, right? So feedback decreased. The problem with this kind of system, if you did not have this measurement system, is that you couldn't actually know what's happening. But because we had the measurement system in place, these the drop in C set was detected, right? Because we were getting negative feedback from customers. We could actually look into the tracing decisions and see that the agent was looking at a policy policy document that was outdated. So the the new policy document was not updated in the vector database. The embeddings did not come through.
然后到了第七、第八周,我们才开始谈模型。这时候我们已经有了评测数据集,就能拿不同的模型在这个数据集上跑,看它们的回答,跟期望回答做比对,算出一个准确率的数字对吧?这帮我们决定该用哪个模型。这个决策没花多少时间对吧。就像我在开头讲的,我们曾经花好几周争论该用哪个模型,但换成这另一套打法之后,你其实可以非常快地搞定。一旦这个搞定,我们就把前面讲的那些围绕可观测性、评测、各层评测的东西全都拼起来。等我们有了这套能让 AI 变得可见、可衡量、可问责的系统,我们才开始把它推上生产。结果就是这样——上线六周后,我们算了运维指标,当然包括准确率、分流率、响应时间、客户 CSAT。但这里更重要的是:上线几周后出了个问题——其中一件事是,这家银行改了一些跟利率相关的政策。改政策的时候,他们确实给客户发了邮件、或者在手机银行 App 里发了政策变更的通知。但当客户跑到聊天机器人上追问后续问题时,机器人给不出正确答案,他们就给这些回答点了踩。于是我们收到了这些反馈对吧,反馈(评分)就下降了。这类系统的麻烦在于,如果你没有这套衡量系统,你根本不知道发生了什么。但因为我们这套衡量系统在位,CSAT 的这次下滑被检测到了对吧——因为我们一直在收到客户的负面反馈。我们就能去看那些追踪到的决策,发现 agent 当时引用的是一份已经过期的政策文档。也就是说,新的政策文档没有更新进向量数据库,embedding 没有生成进去。
[30:34] Sandipan
Because it did not come through, it was giving it stale answers. And that's when we went and fixed that. But it it was all possible because we built that those systems that that led us to to detect this. Before you go, I I generally in these sessions I share different artifacts that you can take away. I have a QR code at the end for you to download and you will find multiple artifacts. One of the important artifacts that I want to talk about is the production incident playbook. This is something that a lot of us tend to miss when we work in AI projects. And this playbook is basically a definition of what needs to happen when things fail in production. First, you detect using your eval dashboard. Then you diagnose using your tracing as I explained. Then you contain. So basically, you you are versioning your prompts. Is there a Is there is a If there is a problem with the prompt, you you take that prompt out, right? And start the changes, deflect it to a human. Or in my multi-agent orchestration video, I've talked about multiple fault tolerance failure recovery patterns around saga pattern, compensation pattern, and circuit breaker pattern that you can look into the video. I've explained them in details on how you can handle them. And then you use the test case library to fix. So you look into LLM's judge reports, you look into your evaluation data set reports. Then you fix your problem. Once you fix your problem, you put those test cases in your data set, right? And and create that eval suite that is a living system that will keep growing. And and and you you keep improving your AI system based on that, right? But this playbook needs to be in place. When it runs in production, you will need to integrate it with your ITSM system so that it alerts the right person at the right time. You know, a lot of these organizations would have existing ITSM systems, right? So which which is used for alerting and you know, making sure that the downstream systems don't get affected, etc.
因为没生成进去,它就一直在给过时的答案。我们这才去把它修好。但这一切之所以能做到,全靠我们搭好了那些系统,让我们能检测到这个问题。在结束之前——我在这类分享里通常都会给一些可以带走的资料。我在最后放了个二维码,你们扫码就能下载,里面会有好几样资料。其中我想特别讲的一样,是「生产事故应对手册」(production incident playbook)。这是我们很多人做 AI 项目时容易漏掉的东西。这份手册本质上就是定义清楚:当生产环境出故障时,该按什么流程走。首先,你用评测仪表盘来检测。然后,用我前面讲的追踪来做诊断。接着是控制(contain)。本质上就是你在给 prompt 做版本管理——如果某个 prompt 有问题,你就把那个 prompt 撤下来对吧,开始做改动,把它转给人来处理。或者在我那期多 agent 编排视频里,我讲了好几种容错、故障恢复模式,比如 saga 模式、补偿(compensation)模式、断路器(circuit breaker)模式,你可以去视频里看,我详细讲了怎么处理它们。然后你用测试用例库来修复。你去看 LLM as judge 的报告,看你评测数据集的报告,然后把问题修掉。修好之后,把那些测试用例放进你的数据集里对吧,组成那套不断生长的评测套件,再据此持续改进你的 AI 系统。但这份手册必须就位。当它在生产里跑起来时,你需要把它跟你的 ITSM 系统打通,好在对的时间提醒对的人。很多这种组织都已经有现成的 ITSM 系统对吧,用来做告警、确保下游系统不受影响等等。
[32:38] Sandipan
So once once you have this in place, you can go and stitch it together to other systems. So what can you do tomorrow, right? Start with If If you have a project in mind, start with defining success. Success not from the technical sense, from the business sense. What it means means for the business, right? Come up with a few examples of what good answers look like. And and create a data data set of that. And then build that pipeline using simple Python code. See if you can automate that so that when you run AI and get some response, it can become it can go and compare You can go and compare the answer against that data set and then that can be delivered to to the to the customer. Now, these are three lessons that I have learned while doing these things with my you you know, my easily miss. The test case library, as I explained, is a growing system. It will grow over time. And because it grows over time, you need some sort of governance around it. You need a owner, right? You need to You need to figure out which test cases relate to what kind of problem. So that whenever you go back to it, you can you can relate your answers to those sort of problems, right? If it is a security If it is login, so you can say that the agent did not ask for login credentials when the customer asked the answer. And all those kind of issues can be put under a security category within that data set. So to categorize the rows in your data set so that you can pick up what changed and compare it with them. The second is prompt versioning. Now, when you start versioning prompts using Git, you know, we all know when you put Git message commit messages tend to be simple commit messages. But you have to put governance around what kind of commit messages you are putting in when you're changing these prompts because you need to understand when a prompt was changed, for exact what reason it was changed, right? What was the failure that caused this prompt to be changed? What kind of failure would it address and what would it correct, right? In the next version. That needs to be documented.
所以一旦这个就位,你就能把它跟其他系统拼接起来。那么你明天可以做点什么呢?如果你心里已经有个项目,那就从「定义成功」开始。这个成功不是技术意义上的,而是业务意义上的——它对业务到底意味着什么对吧。给出几个「好答案长什么样」的例子,把它们做成一个数据集。然后用简单的 Python 代码搭起那条流水线,看你能不能把它自动化,让你跑完 AI、拿到某个回答后,它能自动去拿这个答案跟那个数据集做比对,然后再把答案交付给客户。下面是我做这些事情过程中学到的三条经验,都是很容易被忽略的。第一,测试用例库,就像我讲的,是个会生长的系统,它会随时间不断变大。正因为它会不断变大,你得围着它做一些治理。你需要一个负责人对吧,你得搞清楚哪些测试用例对应哪类问题,这样每次回头看时,你都能把你的答案跟那类问题对应起来对吧。如果是安全类、比如登录相关,你就能说:客户问问题时,agent 没有要求登录凭证。所有这类问题都可以归到数据集里的「安全」类目下。也就是给数据集里的每一行分类,这样你才能拎出来「改了什么」,再拿它们去对比。第二是 prompt 版本管理。当你开始用 Git 给 prompt 做版本管理时,我们都知道,写 Git commit message 时大家往往随手写得很简单。但你必须围绕「改这些 prompt 时该写什么样的 commit message」做治理,因为你需要搞清楚:一个 prompt 是什么时候被改的、到底为什么改对吧、是什么样的失败导致它被改的、它要解决的是哪类失败、在下一个版本里它要纠正什么对吧。这些都得记录下来。
[34:45] Sandipan
Otherwise, it becomes difficult because when you go back and look into prompt versioning and look at different versions and you cannot trace why why those changes were made, then it becomes difficult to track what's happening. The third, the layer three evals, right? So the behavioral evals that I was talking about around tool calls and stuff like that, they can be really expensive as you grow your eval data set as well. So when you have a wrong tool call, for example, and you want to correct that system, when you correct it and run it against the eval data set, you have to basically run it against let's say if you have got 300, 400, 500 rows in the data set, you have to run it against them. And you do all the testing again and again and again and again and again, that can cost you a lot of money. So you have to put some governance around that. So for example, when in your continuous integration pipeline, when you do the prompt change, you can actually put some checks around just just selecting a small subset of the eval data set to do the testing. And you only do the full test when you merge to the main branch. So you can put these kind of decisions in place so that you can reduce cost around around you know, expensive eval decision. If you scan this QR code, it'll take you to a Google Drive link where I have put some examples on some of these how these templates look like, what evaluation checklist should look like. I've given you some guide on set setting up tracing with open source technologies so that you can quickly set up some tracing and start testing in the test environment before you decide on what kind of tools you want to use. Thank you very much for listening to me. This is This QR code will take you to my LinkedIn profile. So I share I have a newsletter where I share this kind of topics every week. So if you're interested, you can join. It's free. I basically share what I learn in in the field working with customers, right? So it might be useful for you.
否则就会很麻烦,因为当你回头去看 prompt 的版本管理、看不同版本时,你却追不出这些改动当初为什么做,那要追踪到底发生了什么就很难了。第三,是第三层评测对吧。就是我前面讲的那种围绕工具调用之类的行为评测(behavioral evals),随着你的评测数据集变大,它们的成本会变得相当高。比如你有一次错误的工具调用,你想去纠正这套系统,纠正完之后再拿它去跑评测数据集——假设你数据集里有 300、400、500 行,你就得拿它一行行全跑一遍。你一遍又一遍、反反复复地做这些测试,那会花掉你一大笔钱。所以你得围绕这个做一些治理。比如在你的持续集成流水线里,做 prompt 改动时,你其实可以加一些检查:只挑评测数据集里的一小部分子集来跑测试,只有在合并到主分支时才跑全量测试。你可以把这类决策固化下来,这样就能在那些昂贵的评测决策上省下成本。如果你扫这个二维码,它会带你到一个 Google Drive 链接,我在里面放了一些示例,展示这些模板长什么样、评测清单该是什么样的。我还给了你一些用开源技术搭建追踪的指引,这样你可以先快速搭起一些追踪、在测试环境里开测,再去决定你想用什么样的工具。非常感谢大家听我讲。这个二维码会带你到我的 LinkedIn 主页。我有一份 newsletter,每周分享这类话题,如果你感兴趣,可以加入,是免费的。我基本上就是把我在一线跟客户打交道时学到的东西分享出来对吧,也许对你有用。
[36:46] Sandipan
Thank you very much.
非常感谢大家。
[36:47]
[applause]
[掌声]