ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.92 · 全文

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation

频道: AI Engineer
视频: https://www.youtube.com/watch?v=jt1Pbr_n6oU
原文语言: en
统计: 共 50 轮


[0:01]

[music] Yes. So my talk today is about the title your data models remote. We have a enterprisewide platform that we had just rolled out here this past month. And so I'll go into details on this. I'll give you some hopefully some practical lessons here and why we made decisions we made for this uh how how you could picture your processes within a similar type framework.

好的。我今天要讲的主题就是这次演讲的标题——你的 data model 就是你的 moat。我们有一个企业级的平台,上个月刚刚在这里上线。我会详细讲讲这个平台,希望能给大家分享一些实用的经验,讲讲我们当时为什么做出这些决策,以及你可以怎样把自己的流程套进一个类似的框架里去思考。


[0:35]

So first just a quick introduction. So this gets into the the title here the talk and the the framing of you know what I hope you take from this but with AI moving very fast at the frontier what what's defensible you know you can move you can build things very quickly with clogged code um but once you push things to production there's constraints you find how much of your your uh deployed stack do you want to actually own you

首先简单做个介绍。这就引出了这次演讲的标题,以及我希望大家能从中获得的思考框架。在 AI 前沿发展如此之快的今天,到底什么才是有防御力的?你可以用 Claude Code 非常快速地把东西搭起来,但一旦把它推到生产环境,你就会发现各种约束——你到底想真正拥有自己所部署的这套技术栈中的多少部分?


[1:04]

know there's monitoring there's upkeep there's uh you know people building dependencies off your stack that you have to be prepared to handle. How much uh appetite do users have for decentralized access? This gets into I'll show you what we built, but you know what's what's the access point for users? You know, is it another chat app?

你要考虑监控、要考虑维护,还有别人会基于你的技术栈去构建各种依赖,这些你都得做好准备去应对。用户对分散式的访问方式又有多大的接受度?这就要说到我们到底构建了什么——用户的访问入口是什么?是又一个聊天应用吗?


[1:23]

Is it clawed? Is it chat GPT? Uh is it something else? What's your product differentiation from from those different SAS products? And so our team then you know with this context in mind you know thought through here you know what's our skill set here what's our competitive advantage in this environment and this is what I you really hope that you you take from this talk and you picture yourself in this but our moat here was our understanding

是 Claude 吗?是 ChatGPT 吗?还是别的什么?相比那些各式各样的 SaaS 产品,你的产品差异化又在哪里?于是我们团队就带着这样的背景去思考:我们的技能优势在哪里?在这样的环境下我们的竞争优势是什么?这也正是我特别希望大家能从这次演讲中带走的东西,请把你自己代入进来。而我们的 moat,正是源于我们的理解——


[1:51]

of our internal processes the tacet knowledge that you need to to run successful AI and this is true I think no matter how good AI gets how good models get and new releases that different companies put out when when Mythos comes out or when there's a new app from Claude. Yeah, I'm not I'm not worried because the part that we've built is the defensible part that that's that's durable.

——也就是对我们内部流程的理解,那种成功运行 AI 所必需的 tacit knowledge(隐性知识)。我觉得无论 AI 变得多强、模型变得多好,无论各家公司推出什么新版本,无论是 Mythos 发布,还是 Claude 出了新的应用,我都不担心。因为我们所构建的这一部分,才是有防御力的,是能够长久持续的。


[2:15]

So, these are the I I'll tell you what this means here in more detail, but these are the the processes tacet knowledge that we've modeled into what we call the strategic intelligence platform or SIP. and it rolled out here this past month in production for enterprise use across uh the Gates Foundation. So about 4,000 people.

这些就是……稍后我会更详细地讲它到底意味着什么,但这些就是我们建模进所谓“战略情报平台”(Strategic Intelligence Platform,简称 SIP)里的那些流程和 tacit knowledge。它上个月刚刚上线,在生产环境中供整个 Gates Foundation 的企业内部使用,大概涉及 4000 人。


[2:36]

So first I know this is an engineering talk but the the the scope of this talk gets into data modeling internal operations processes and so I want to give very quick background here over what the Gates Foundation does because then this is what we're we're modeling.

首先,我知道这是一场偏工程的演讲,但这次分享会涉及 data modeling、内部运营流程这些内容,所以我想先非常快速地介绍一下 Gates Foundation 是做什么的,因为这正是我们所要建模的对象。


[2:52]

So, as you're probably familiar, the the Gates Foundation has a has a very wide scope and it's a very ambitious work that we've been doing for the past 25 plus years. And there's all kinds of, you know, broad initiatives that we're doing, you know, whether it's for uh child mortality, whether it's for nutrition, agriculture, uh education.

大家可能都比较熟悉,Gates Foundation 的业务范围非常广,我们过去 25 年多来做的都是很有雄心的工作。我们有各种各样的大型项目,比如降低儿童死亡率,比如营养、农业、教育等等。


[3:14]

And these are kind of broadly the the different buckets that these different initiatives fit fit into. creating market incentives, spurring an innovation, collaboration between public and private sectors. And then the fourth one here kind of gets into the the lens that we're building here. You know, high quality data trying to derive datadriven insights from the from the actual investments, the the grants that we've

这些项目大致可以归入几个不同的类别:创造市场激励、激发创新、推动公共部门与私营部门之间的合作。而第四个类别,正好切入到我们正在构建的这个视角——用高质量的数据,试图从我们实际的投资、也就是我们发放出去的 grant 中,提炼出数据驱动的洞见。


[3:38]

put out. And over 25 years, there's a ton of structure. There's a ton of data that's developed. And trying to extract those insights at scale is difficult. And that's what we're trying to solve. So this slide here is a snapshot of the of some of the different uh of the work that went out in 2023 within the foundation. This gives you an idea. I just put this here to to show some of the structured the structure that we

在过去 25 年里,积累了大量的结构,也产生了海量的数据。而要大规模地从中提取这些洞见是很困难的,这正是我们想要解决的问题。这张幻灯片是基金会 2023 年发放的部分工作的一个快照,给大家一个直观的概念。我放这个,是为了展示我们所面对的一些结构化的东西——那种结构,


[4:06]

have that we're working across. So you have over 2,000 grants in one year. Many of these are 5 million plus uh many 100 plus countries that uh that that are targeted with these grants uh alumni. So there's 4,000 different employees of the foundation. Um you know many different strategies within the foundation the US within the US across almost all the states grantees the total annual dispersement over 7 billion dollars.

——就是我们要横跨去处理的结构。比如一年就有超过 2000 笔 grant,其中很多都在 500 万美元以上;这些 grant 覆盖 100 多个国家。基金会有 4000 名员工,内部有许多不同的战略方向;在美国境内几乎覆盖了所有的州;受资助方众多,每年的总拨款超过 70 亿美元。


[4:37]

And so this gives you some idea of structure that we're we're working with. And this one just finally here when I show the data model this will make more sense. But we have different divisions that that funding goes out through different divisions. And so this breaks down some of those divisions. So you can see different priorities and it'll make more sense in a second here. But global development, global health, uh gender

所以这能让大家对我们所面对的结构有个概念。最后这一点,等我展示 data model 的时候会更好理解。我们有不同的部门(division),资金是通过不同的部门发放出去的。这里就拆解了其中一些部门,你可以看到不同的优先事项,稍后就会更清楚了。比如 global development(全球发展)、global health(全球健康)、gender


[4:59]

equality, USP are just a sample of the different divisions. Okay. So the the fun stuff here now I hope the uh strategic intelligence platform so in a in a in a nutshell here structuring operational data for agentic retrieval. So we're building a knowledge graph with the idea of the agent consumer and here is an endto-end look of what this looks like. So we we have different systems of record structured unstructured these have been siloed traditionally

equality(性别平等)、USP,这些只是众多部门里的一小部分示例。好,现在到了有意思的部分了——战略情报平台。一句话概括,它就是把运营数据结构化,以支持 agentic retrieval(面向 agent 的检索)。所以我们在构建一个 knowledge graph,核心理念就是把 agent 当作数据的消费者。这里是它端到端的整体样貌。我们有不同的 systems of record(记录系统),有结构化的、也有非结构化的,这些数据传统上一直是彼此孤立的。


[5:36]

the so part of our team here the work has been to create what's essentially a data lakehouse putting everything under one roof. This is our internal enterprisewide data. It's also different different programmatic data that are uh outputs of different investments.

所以我们团队的一部分工作,本质上就是构建一个 data lakehouse,把所有东西都放到同一个屋檐下。这既包括我们企业内部全域的数据,也包括各种项目性的数据,也就是不同投资所产生的产出。


[5:54]

Once it's there it's easy for us to consume. So we have a data curation layer that does different processing to it and then finally SIP here at the end with aentic chat agentic workflow as the as the the UX you how users are consuming our platform and so it's a cross system semantic graph layer that agents can res across okay uh so some of this I'll try to speed through here just for the sake of time but this one is critical

数据一旦进来,我们就很容易去消费它。我们有一个 data curation(数据治理)层,对数据做各种处理,最后就到了 SIP。它以 agentic chat、agentic workflow 作为 UX,也就是用户消费我们平台的方式。所以它是一个跨系统的语义图层(semantic graph layer),agent 可以在它之上进行推理。好,为了节省时间,有些内容我会讲快一点,但这一点非常关键。


[6:20]

when you're dealing with systems of record with lots of complexity engagement ment is critical. This is something that we've we've found here repeatedly. We have to engage data owners to understand, you know, this tacet knowledge we're trying to to model. What's the full meaning of different fields, the structure of the data set, how do we join things together? How do we uh understand limitations, systematics of the data,

当你处理这些极其复杂的 systems of record 时,沟通协作(engagement)至关重要。这是我们反复体会到的一点。我们必须去和数据的负责人(data owner)沟通,才能理解我们想要建模的这些 tacit knowledge:不同字段的完整含义是什么、数据集的结构是怎样的、我们该如何把各部分数据 join 到一起?又该如何理解数据的局限性、数据中的系统性问题,


[6:41]

safeguards, security trimmings, uh reporting conventions? You know, it's not enough just to answer it a question a certain way. You have to answer it the way that it's been answered in the past. And so, this is the comes back to the moat here. This is the procedural understanding tacet knowledge that AI needs and it's yeah it's the part that we that we own that's you know that's ours that um and that's what we're modeling here.

还有各种安全保障、security trimming(安全权限过滤)、报告惯例?要知道,仅仅用某种方式把一个问题回答出来是不够的,你必须按照过去一贯的方式来回答它。这又回到了 moat 这个话题上。这就是 AI 所需要的那种流程化的理解、tacit knowledge,而这正是我们所拥有的、属于我们自己的东西,也正是我们在这里所建模的内容。


[7:09]

Okay. So going back here just very quickly for this one this is the a snapshot here of different data curation considerations that we're that go into this pipeline. So you have [sighs] for different data sets whether it's structured unstructured there's different pre-processing filtering dduplication there's an order to different documents there can be uh inconsistencies across documents those need to be uh handled up front there's

好,我们快速回到这一页。这是我们这条 pipeline 中所涉及的各种数据治理(data curation)考量的一个快照。针对不同的数据集,无论是结构化还是非结构化的,都有不同的预处理、过滤、去重(deduplication);不同文档之间是有先后顺序的,文档彼此之间也可能存在不一致,这些都需要在前期就处理好;还有……


[7:31]

extraction so structured field extraction semantic chunking for unstructured documents if you have figures you need to convert this into text in some way so you can do retrieval across this uh various forms of tagging that these can form connections in your graph structure extra metadata that you create during this pipeline. Then that becomes different properties in your graph. And then the third bucket here,

……抽取。比如结构化的字段抽取、针对非结构化文档的语义分块(semantic chunking);如果文档里有图表,你需要用某种方式把它转成文本,这样才能对它做检索;还有各种形式的打标签(tagging),这些标签可以在你的 graph 结构里形成连接;以及在这条 pipeline 过程中生成的额外 metadata,这些之后就会变成你 graph 里的各种属性(property)。然后是第三大类,


[7:54]

governance. This is a important one that I think AI makes more acute things that were that were accessible previously. They're much more accessible now with with AI. And so you have to consider this. Your risk sphere is is larger. So things like PII need to be masked. you need to reconsider different uh sensitive data classifying this um making sure that there's the right entitlements for each user who's accessing your system.

治理(governance)。这一点很重要。我认为 AI 让一些原本就能获取到的东西变得更加尖锐了——有了 AI,这些数据现在变得容易获取得多。所以你必须考虑这一点,你的风险面变大了。像 PII(个人身份信息)这样的数据就需要做脱敏;你需要重新审视各种敏感数据、对它们进行分类;确保每一个访问你系统的用户都拥有恰当的权限(entitlement)。


[8:23]

Okay, so that's the overview here. The this the the data model itself. Now, this is the part I'll walk through here. There's a a nice animation here, but hopefully the takeaway is you can picture your own your own organization story within what I show here. I'll get somewhat technical but it's only to hope hope hopefully to give you an idea of how we how we solved our problem and then you can hopefully uh model this to

好,这就是整体的概览。接下来是 data model 本身,这部分我会带着大家走一遍。这里有个很不错的动画,但希望大家能带走的关键是:你可以把你自己组织的故事,代入到我接下来展示的内容里去。我会讲得稍微技术性一些,但这只是为了让大家了解我们是怎么解决自己的问题的,然后希望你也能把它套用到


[8:46]

yours as well. Graph is very flexible um practical representation of a physical model. [snorts] Okay. So I'll zoom through a few of these here but the just the the entry point here we have over 80 different strategy teams. These teams have annual reviews that happen. This is how the budgeting for each year is derived. And then so we model this here in the graph. The the the meetings are where unstructured documents uh enter

……你自己的场景里。Graph 非常灵活,是对物理模型的一种很实用的表示方式。好,我快速过一下这几页。先看这个入口点——我们有超过 80 个不同的战略团队(strategy team),这些团队每年都会做年度评审,每一年的预算就是这样得出来的。于是我们把这个过程建模到 graph 里。这些会议正是非结构化文档进入这个系统的入口,


[9:12]

into this system from but then there's they have a structured connection to your other systems of record. Um what I show here is a conceptual data model. So it's flat. So you're not seeing the instantiation. The actual graph there's you know many different nodes. Cardonality is it one one to n. So the actual graph it's you know even more complicated. But for the data model itself, let's let me let me show you the first different

……但同时它们又与你其它的 systems of record 之间保持着结构化的连接。我这里展示的是一个概念性的 data model,所以它是扁平的,你看不到具体的实例化。真实的 graph 里有非常多不同的节点(node),基数(cardinality)是一对多(one-to-n)的,所以实际的 graph 要复杂得多。但就 data model 本身而言,我先给大家展示第一种不同的


[9:37]

hier. So we have multiple hierarchies that exist within what we've modeled the there's different types of hierarchies you can have. In this case, this is a hopefully you can see all this very well, but it's a it's an it's a um additive DAG. So there's a all five levels here of this hierarchy from the top to the bottom matter.

……层级(hierarchy)。在我们所建模的内容里,存在着多个层级,而层级也可以有不同的类型。在这个例子里——希望大家都能看清楚——它是一个可累加的 DAG(有向无环图)。所以这个层级从上到下的全部五层都是有意义的。


[9:59]

So you have to consider everything together. And so then there's different rollup patterns you can do to work across this this sort of pattern. In our case, we have a in path shortcut here that connects the funding path. Uh funds to bow is where we have the um the budget for each of these different funding teams that that's stored.

所以你必须把所有层级放在一起考量。因此你可以用不同的 rollup(汇总)模式来贯穿处理这一类结构。在我们的例子里,这里有一个 in-path(路径内)的快捷连接,把资金路径连了起来。funds-to-BoW 就是存放每个不同资金团队预算的地方。


[10:24]

So we have funding what's the the internal funding teams have portfolios. These portfolios then go towards different investments. Multiple funding teams fund an individual investment. So it's a endtoend relationship there. The investments are the thing that are our product. It's our it's our our business.

我们有资金……内部的资金团队(funding team)拥有各自的投资组合(portfolio),这些 portfolio 又会流向不同的投资(investment)。多个资金团队会共同资助同一笔投资,所以这里是一种端到端的多对多关系。而这些投资才是我们真正的产品,是我们的业务所在。


[10:44]

But internally we have funds that then prioritize different different types of investments. That's what's shown here. And so you can take this down to the transaction level or you can have different uh annualbased aggregations that you map here as well.

但在内部,我们有各种资金(fund),它们会对不同类型的投资划分优先级,这里展示的就是这个。你可以把它一直细化到交易(transaction)层面,也可以在这里映射出各种基于年度的聚合数据。


[10:59]

[snorts] >> And then from investment, there's a lot of interesting things you can do. You can map to all the different organizations and you can have different types of organizations and there's actually a lot here that is still kind of green space that we want to fill in.

然后从投资(investment)出发,你可以做很多有意思的事情。你可以映射到各种不同的组织(organization),而且组织也可以有不同的类型。实际上这里还有很大一块仍然是待开发的空白地带,是我们想要去填补的。


[11:11]

We have all these different observables that people have produced in the investments that we want to model here. So publish reports, products, you know, all this stuff is structured and connects to the entire uh organizational picture. >> [snorts] >> So I mentioned that there's different hierarchies. This is the second type of hierarchy. At this hierarchy, each level matters in and of itself. And so it's not a a DAG necessarily. And so you can

在这些投资中,人们产出了各种各样可观测的成果,我们也想把它们建模进来,比如发表的报告、产品等等,所有这些都是结构化的,并且连接到整个组织的全景图上。前面我提到过有不同类型的层级,这是第二种层级。在这种层级里,每一层本身都独立地有其意义,所以它不一定是一个 DAG。因此你可以


[11:37]

actually do things like precomputing the the some of these these different shortcuts. So the hierarchy it goes from the top to the bottom contains connects. It this is showing the investment management side of the of the organization. And there's concepts of direct team management. So one team at like team level two manages the investment. But then there's also a concept of indirect management. So the uh children below

……真的去做一些事情,比如预先计算其中一些不同的快捷连接。这个层级从上到下是“包含(contains)、连接(connects)”的关系。这里展示的是组织中投资管理的这一侧。这里有“团队直接管理”的概念,比如某个处于第二层级的团队直接管理某笔投资;但同时也有“间接管理”的概念,也就是说它下面的子级……


[12:05]

team level two still should be attributed to the team level two. And so there's different things you can different games you can play with these sort of rollups to precomputee. I don't know if you can see this, but rollup manages m is a is a a derived edge that we that we create after we create the contains and manage manages edge.

第二层团队的东西,依然应该归属到第二层团队上。所以针对这类 rollup 预计算,你可以有各种不同的玩法。不知道大家能不能看清,这个 rollup manages 是一条派生边(derived edge),是我们在创建了 contains 边和 manages 边之后,再额外生成出来的。


[12:27]

So I've shown two different two different lenses for one investment. There's the funding lens, the management lens and you can model both of these here. Then within within the graph, a third hierarchy here is people. You have organizations, you have org charts,

所以针对同一笔投资,我展示了两个不同的视角(lens):一个是 funding 视角,一个是 management 视角,这两者你都可以在这里建模。然后在这个 graph 里,第三个层级是「人」。你有组织,有组织架构图(org chart),


[12:43]

and you have people who are owners, you have people who are attendees of meetings, you have people who are uh directors. There's all kinds of different roles they have. You can model these here. You can have their their uh you know, who they report to, what their uh team structure is. And these all are structured data that connects across systems. Traditionally, they existed in a just a HR source system, but they're

有的人是 owner,有的人是会议的参会者,有的人是 director。他们有各种各样不同的角色,这些你都可以在这里建模。你可以记录他们向谁汇报、他们的团队结构是怎样的。而这些全都是能够跨系统连接起来的结构化数据。传统上,它们只存在于 HR 这一个源系统里,但它们


[13:04]

relevant for the context of the the full story. And then that leads to this connectedness. So we have different source systems that were siloed. We to understand the entire picture for the agent to understand correctly across the structure you need to find these common B these common uh these common uh entities that you stitch together. And so that's what's shown here. These are different source systems but they're related quantity entities that exist

对于理解整个故事的上下文其实非常相关。这就带来了这种连接性。我们原本有很多彼此孤立(siloed)的源系统。要理解全貌、要让 agent 能够正确地跨这个结构去理解,你就得找到这些共同的 entity,把它们缝合(stitch)在一起。这里展示的就是这个。这些是不同的源系统,但它们之间存在着相互关联的 entity。


[13:36]

there. And now the agent can traverse here and understand this pretty complicated organ organizational structure. One last part here that I haven't shown yet is the the document part. So, and this is still there's there's a lot more we can do to this part. We've just been uh ingesting one different document source so far, but this is where you combine unstructured and structured. And this gets into part of the magic here

现在 agent 就可以在这里做 traversal,理解这个相当复杂的组织结构。这里还有最后一部分我还没展示,就是 document 这一块。这部分其实我们还有很多可以做的,到目前为止我们只接入(ingest)了一个文档源,但这正是你把非结构化数据和结构化数据结合起来的地方。这也涉及到这里的一部分魔力,


[14:05]

that you can model with Neo4j. But we have meetings that have documents. Documents then have different semantic sections that you can or chunks that you can uh that you can model here. You can put full text indexes across these to to aid in the different uh search and retrieval approaches for the agent. There could also just be a pure graph retrieval that that the agent does. And then all these things then connect back to your your your main

也就是你可以用 Neo4j 来建模的东西。我们有会议,会议里有文档,文档又有不同的语义段落(semantic section),或者说 chunk,这些你都可以在这里建模。你可以在这些之上建立全文索引(full text index),来辅助 agent 各种不同的搜索和检索方式。也可以是 agent 做的纯 graph 检索。然后所有这些东西又都连回到你的主


[14:33]

organizational structure. So then as a whole this is what the data model looks like. So I've been zooming in here now. You can see the the full interconnectedness of this four different systems one graph uh one semantic layer that's exposed through an MCP then to the to the agents.

组织结构上。所以整体来看,这就是这个 data model 的样子。我之前一直在放大看局部,现在你能看到整体的相互连接:四个不同的系统、一个 graph、一个语义层(semantic layer),通过 MCP 暴露出去,再交给 agent。


[14:56]

And so this is the so if you think of the agent's perspective, this is the the structure that it can dynamically discover and reason across at query time. And for the developer, it's also a very cool thing because it exposes, you know, what you don't know about your your the thing you're modeling. You very soon you find out that there's a gap in your understanding or there's some data set that you're not, you know, fully

所以,如果你站在 agent 的视角看,这就是它在查询时(query time)能够动态发现、并在其上进行推理的结构。而对开发者来说,这也是件很酷的事,因为它会把你对自己建模对象其实并不了解的地方暴露出来。你很快就会发现自己的理解里存在缺口,或者有某些数据集你还没有完全


[15:18]

including. And so this this process in in of it in and of itself is very valuable. Okay, let me give you a sense here what we do with this now. So this is the I showed you the platform, the graph, but then how does this relate to AI? So we've we've connected this through MCP and I you know I discussed earlier what the what's durable, what's defensible to us. What was not defensible was the was the the chat interface was the UI and

纳入进来。所以这个过程本身就非常有价值。好,接下来我给大家讲讲我们现在拿这个来做什么。我给你们展示了平台、展示了 graph,那这跟 AI 又有什么关系呢?我们是通过 MCP 把它连接起来的。我前面也讨论过,对我们来说什么是持久的、什么是有护城河的(defensible)。没有护城河的,是 chat 界面、是 UI,


[15:49]

even in some cases the the general chat cases the um you know the agent interaction and so we users themselves are included already or chat GPT and so we serve the platform where they are and so it's served here now through MCP here's an example just kind of a innocuous uh question here but Neoforj has some off-the-shelf uh MCP servers Here we've actually modified these quite a bit here. We forked it. And then there's various updates to the schema.

甚至在某些情况下,连通用的 chat、连 agent 交互都算不上护城河。所以我们要到用户所在的地方去服务他们——用户本身可能已经在用某个界面,或者在用 ChatGPT。所以我们就把平台服务到他们所在的地方,现在就是通过 MCP 来提供。这里有个例子,是个挺普通、无关紧要的问题。Neo4j 有一些现成的(off-the-shelf)MCP server,我们其实在这基础上做了相当多的修改,我们 fork 了它,然后对 schema 做了各种更新。


[16:19]

Uh things to to pass state back to the to to our system. You know, the conversation uh ids, the the message uh numbers, stuff like this we we've we've modified in these MCP tools. But so that's a general chat experience. That's one entry point. The other part that we're building right now too that's very exciting is more constrained workflow experiences. And so these can also be offered through things like co-work uh clawed chat. And you can do

比如一些把状态传回我们系统的东西——对话的 id、消息的编号,诸如此类,我们都在这些 MCP 工具里做了修改。这就是一个通用的 chat 体验,是其中一个入口。我们现在还在做的另一部分,也非常令人兴奋,就是更受约束的(constrained)工作流体验。这些也可以通过像 co-work、Claude chat 这样的方式来提供。你可以做


[16:49]

things like um you can have your you can have MCP apps be the the you the standard entry way that users access you know different UIs that are uh ported into your your your your chat experience and you can have different sandbox based agents that then run the the workflow.

比如说,你可以让 MCP app 成为用户访问的标准入口,让各种 UI 被移植(port)进你的 chat 体验里;你还可以有各种基于 sandbox 的 agent,来运行这些工作流。


[17:07]

And so these are these are active things that we're working on. It helps to constrain the experience compared to chat, but it at the same time it pulls from that same knowledge graph-based uh back-end platform. Okay, I've got a couple minutes. I'll kind of speed through this, but the the way eval relate to data modeling is that as you're doing eval, you find you find gaps. You find ambiguities in your data model. you find ways in which users are

这些都是我们正在积极推进的事情。相比纯 chat,它有助于把体验约束住,但与此同时,它又是从同一个基于 knowledge graph 的后端平台里取数据的。好,我还有几分钟,我快速过一下。eval 和 data model 的关系在于:当你做 eval 的时候,你会发现缺口,会发现 data model 里的歧义,会发现用户


[17:38]

asking questions that uh that are ambiguous or it's um you not it's not returning things that conform with the reporting standards. So what we've done here then is we've worked with data owners. We've we've uh built targeted uh eval questions that that they that match their reporting standards. We've separated these into different complexity tiers. One challenge is that the the structured data is constantly changing. So we have to have the graph

提问的方式是有歧义的,或者说返回的结果不符合报告标准(reporting standard)。所以我们的做法是,和数据的负责人(data owner)一起合作,构建了有针对性的 eval 问题,让它们匹配他们的报告标准。我们还把这些问题按不同的复杂度分成了不同的层级(complexity tier)。有一个挑战是,结构化数据是在不断变化的,所以我们必须准备好 graph


[18:06]

query itself that we that we create for each of these different questions and then at runtime for the eval we we pull from the live graph and then we compare that to what the agent is delivering for that question and so that's what's shown here then there's a feedback loop that you can do for this. So as you're running an eval pipeline, an eval structure pipeline,

查询本身——我们为每一个不同的问题都写好对应的 graph query,然后在 eval 运行时,我们从实时的(live)graph 里拉取数据,再拿它和 agent 针对那个问题给出的结果做对比。这里展示的就是这个。接下来你还可以做一个反馈闭环(feedback loop)。所以当你在跑一条 eval pipeline、一条结构化的 eval pipeline 时,


[18:26]

you have an LLM as a judge, we've modeled things like pass at one uh stability. So if you ask the same question multiple times, you get the same answer back, you can use LLM as a judge to to to measure this. And then there's a feedback loop here that you you can update then your your data model. you can update your uh your domain rules, your uh schema descriptions to help to help uh fill those those gaps that you that you find.

你会用 LLM as a judge,我们建模了像 pass@1、稳定性这样的指标。也就是说,如果你把同一个问题问很多遍,每次都能得到相同的答案,你就可以用 LLM as a judge 来衡量这一点。然后这里有个反馈闭环,你可以据此更新你的 data model,更新你的领域规则(domain rule)、你的 schema 描述,来帮你填补那些你发现的缺口。


[18:58]

Then after you do this, this is a this is just some some eval reporting here that we uh we show the pass at one and the the stability for our system. So we we've gotten this very very strong. the the questions that we that we end up do missing, it tends to be things that are ambiguous in some way. And so it's not wrong, it's just that it's things that might be right but not what the user intended. So that that's kind of the

做完这些之后——这里只是一些 eval 的报告,我们展示了我们系统的 pass@1 和稳定性。我们已经把这个做得非常非常强了。我们最后还会答错的那些问题,往往是某种程度上带有歧义的。所以它其实并不算错,只是它可能本身是对的,但不是用户想要的那个答案。所以这就是那种


[19:23]

constant struggle that we that we that we're working around 30 seconds here. What's ahead for SIP? So we're continue continuing to fill out our existing uh data from systems of records. So things that fit into our current data model. We want to expand the primary graph to additional enterprisewide data sets. There's a lot of there's a lot of demand for a federated graph experience. So, we have a main enterprise system, but we have

我们一直在应对的持续拉扯。这里还有大概 30 秒。SIP 接下来要做什么?我们会继续把现有的、来自记录系统(system of record)的数据补充完整,也就是那些符合我们当前 data model 的数据。我们还想把主 graph 扩展到更多全公司范围(enterprise-wide)的数据集。大家对一个联邦式(federated)的 graph 体验有很多需求。所以,我们有一个主企业系统,但我们也有


[19:47]

specific teams that have their own data that they want to link to this. And so, we're working on how to do this uh different agentic experiences like I mentioned as well. And that's it. Uh, so yeah, please if you want to ask questions, if there's things that you want to talk about, I'll be out back or you can add me on LinkedIn here and you know, keep the conversation going. Thank you.

一些特定的团队,他们有自己的数据,想要链接到这个系统上。所以我们正在研究怎么做到这一点,同时也在做我前面提到的各种 agentic 体验。就是这些了。所以,如果大家有问题想问、有什么想聊的,我会在后面,或者你也可以在 LinkedIn 上加我,我们继续交流。谢谢大家。


[20:10]

[applause] >> [music]

[掌声] >> [音乐]