ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.99 · 全文

AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j

频道: AI Engineer
视频: https://www.youtube.com/watch?v=kRkcNOsRyYg
原文语言: en
统计: 共 272 轮


[0:12]

Hello everyone. Welcome today to get started on AIE. Um, so right here, uh, there's some steps for getting started. I went over this around 10 minutes ago. Um, but basically our workshop that we're going to take today is driven by a website called Graph Academy. And if you go to that first QR code there to the left, uh,

大家好,欢迎来到 AIE 大会,我们现在开始。呃,这边有几个开始前的准备步骤,我大概十分钟前讲过一遍。基本上,我们今天这个 workshop 是靠一个叫 Graph Academy 的网站来带的。如果你扫左边第一个二维码,呃,


[0:36]

that'll take you there. You have to enroll with your email. Um, and then if you go down to set up your environment, uh, there's a code spaces there with everything set up and you can get that rolling now. It'll take maybe about five minutes or so. Um, and you can grab your credentials then as well. There's um an anthropic API key. Uh if you have your own cloud code key um or your your own subscription, please feel free to use

就能直接到那个网站。你需要用邮箱注册一下。然后往下翻到「设置你的环境」那一节,那里有一个已经配好的 Codespaces,你现在就可以让它跑起来,大概要五分钟左右。同时你也可以在那里拿到你的凭证。里面有一个 Anthropic API key。如果你自己有 Claude Code 的 key,或者你自己有订阅,尽管用你自己的。


[1:01]

that. Otherwise, we provided one for you. Um and then there's another key there for reading from BigQuery tables. So with that in mind, uh we'll go ahead and get started with AI on your lakehouse, um which is about context coming in shapes and not necessarily queries.

没有的话我们也给你准备了一个。另外还有一个 key 是用来读取 BigQuery 表的。那么带着这些,我们就正式开始今天的主题:《AI on Your Lakehouse》,讲的是「上下文是以形状(shapes)出现的,而不一定是以查询(queries)」。


[1:21]

So today your team um will be myself. My name is Zach Blumenfeld. I am an AI research engineer at Neo Forj. Uh we also have Ben Squire over there in the back um who is our senior developer advocate um as well as Ryan here in the center um who is our partner architect helping customers get this stuff up and running. So, as you have questions, um I don't have a mic set up for you, but uh go ahead and raise your hand. I'll sort

今天你们的团队呢,首先是我,我叫 Zach Blumenfeld,我是 Neo4j 的 AI 研究工程师。后面那位是 Ben Squire,他是我们的资深开发者布道师;中间这位是 Ryan,他是我们的合作伙伴架构师,专门帮客户把这套东西真正跑起来。所以大家有问题的话,虽然我没给你们准备麦克风,但请直接举手。我会大概


[1:51]

of I'll take time around every 10-15 minutes during natural breaks. Um and and I can take some questions, but also flag Ben and Ryan as well if you're going through some steps and you're having some trouble setting things up. Um and they can help get you unblocked.

每隔 10 到 15 分钟,在自然的间隙留出时间来。呃,我可以回答一些问题;但如果你在做某些步骤、环境搭建卡住了,也可以直接叫 Ben 和 Ryan,他们能帮你解开卡点。


[2:08]

Um so what we're going to talk about today um is really about when you start using uh a lakehouse right there's sort of two sides obviously there's a warehouse which is your structured data and your tables um and then there's the data lake part which is all of your unstructured documents and oftent times what can happen is you're given these tools like text to SQL and vector search and nowadays we don't really have trouble accessing that data Um, but

呃,我们今天要聊的核心是:当你开始用 lakehouse 的时候,其实明显分成两边——一边是 warehouse(数据仓库),也就是你的结构化数据和表;另一边是 data lake(数据湖)那部分,也就是你所有的非结构化文档。而经常发生的情况是,你手里有 text-to-SQL、向量检索这类工具,如今访问数据本身已经不太难了。但是,


[2:38]

sometimes there are still some challenges around how do you give your agent the right type of context, whether or not they can see all the data in the way that they need um, and really take a slice to answer the right type of question. And so what we've put together today inside of our course is an agnostic data model um that basically creates a graph representation uh both from the structured side your warehouse um but then also provides structure for

有时候还是有一些挑战:你怎么给你的 agent 提供正确类型的上下文?它到底能不能以它需要的方式看到全部数据?以及能不能切出恰当的一个切片来回答对应类型的问题?所以我们今天在课程里准备了一套与具体平台无关(agnostic)的数据模型,它基本上会构建出一个 graph 表示——既来自结构化那一侧的 warehouse,同时也为你的文档提供结构,


[3:05]

some of your documents um and then allows you to do a lot of useful stuff with that. Uh we do have a scenario here that we're going to go over. Um, we've created a fictional autofix group, uh, which is, think of it like a a Pep Boys Auto or something. It's a national auto repair chain. And they have all of these bays. They have these libraries of manuals on vehicles, um, as well as safety bulletins and recalls and then a

然后让你能基于它做很多有用的事情。呃,我们准备了一个场景要一起过一遍。我们虚构了一家 AutoFix Group,你可以把它想成类似 Pep Boys 那种连锁汽修,是一家全国性的汽车维修连锁店。他们有很多维修工位(bay),有整套的车辆手册资料库,还有安全公告、召回通知,另外还有一个


[3:33]

warehouse of all of the repairs that they've logged. Um, and you have sort of your floor technicians, right? These are the Danny's that are listed here where they have the cars inside of the bay and they're going to need to ask some questions. Um, you have leadership at this organization that wants to create a co-pilot uh to be able to assist these technicians on the floor. Um, and then obviously you have people like Sam who

记录了他们所有维修工单的 warehouse。然后你有车间一线的技师,对吧,就是这里列出来的这些 Danny 们,他们把车开进维修工位,就会需要问一些问题。这家公司的管理层想做一个 co-pilot,来辅助车间一线的这些技师。然后当然还有像 Sam 这样的人,


[3:58]

are us, the AI engineers who actually have to build the thing. Um, and so if they already have their data in a warehouse, what we're going to be looking at today is the example on BigQuery. So imagine, right, you have your documents inside of cloud storage and then um you have BigQuery as your uh data warehouse, but these patterns are also extendable to data bricks as well as snowflake. Um and essentially if they just create a co-pilot on top of uh

也就是我们——真正得把东西做出来的 AI 工程师。所以,如果他们的数据已经在 warehouse 里了,我们今天要看的例子是基于 BigQuery 的。你可以想象:你的文档放在云存储里,然后用 BigQuery 作为数据仓库;不过这些模式同样可以扩展到 Databricks 和 Snowflake。基本上,如果他们只是在 BigQuery 和已有数据之上直接做一个 co-pilot,


[4:29]

BigQuery and the data that they have um it can pull the data correctly um but sometimes it can be confidently wrong. Um, basically it could get it could pull stuff with vector search from documents and it can do text to SQL um on just a few tables. But when those tables become massive and you get hundreds of tables or when you have large document stores where you have hundreds and thousands and millions of documents um things can

它确实能把数据拉对,但有时候会「自信地给出错误答案」。基本上,它可以用向量检索从文档里捞东西,也可以在少数几张表上做 text-to-SQL。但当那些表变得巨大、你有上百张表的时候,或者当你有一个庞大的文档库、里面有成百上千甚至上百万份文档的时候,东西


[4:56]

get lost and fall through the cracks. And really where graph can come in to help with these shapes um is not only in these single questions that someone might have about how you know I have a broken part that I need to replace and how do I repair this vehicle but oftent times it's going to be on these estate level questions. So things like for example, what are we missing? Like what documentation maybe don't we not have to

就会丢失、就会从缝隙里漏掉。而 graph 真正能靠这些「形状」帮上忙的地方,不只是有人问的那种单点问题——比如「我有个零件坏了要换,这辆车怎么修」——更多时候是那些「全局(estate)层面」的问题。比如说:我们缺了什么?我们是不是缺少某些文档,


[5:22]

cover all of the different cars that are coming in or what documentation maybe are we not leveraging at all. So this is sort of like proving a negative which can be very hard with something like semantic search uh which can only match similar things, right? It can't really find a negative example. Um there's also questions that we have where we want patterns across everything. So say, you know, you might want to ask, well, you

没法覆盖开进来的所有车型?或者,哪些文档我们其实根本没在用?这有点像是在「证伪」,而这件事用语义检索之类的手段是非常难做的——语义检索只能匹配相似的东西,对吧,它没法真的找出「不存在的例子」。还有一类问题是我们想看全量的模式(pattern)。比如说,你可能想问:呃,


[5:46]

know, what are, you know, the common types of patterns that we see? Is there anything that we're failing to fix over and over again? Or, you know, are there, you know, specific groupings of different types of recalls that are popping up and and things of this nature where you really need to traverse the entire data set. Um, and then the other one is just asking how records relate inside of a large SQL store. Um, and

我们最常见的问题模式是哪些?有没有什么东西我们反复修都修不好?或者,有没有哪些特定类型的召回在成组地冒出来?诸如此类——这类问题你必须遍历整个数据集才行。呃,还有一类就是:在一个很大的 SQL 存储里,问这些记录之间是怎么关联的。呃,


[6:09]

when you have one big lookalike schema where you have lots of similar tables, um, how do you understand how to join those together correctly? And so there's going to be three concrete shapes that we'll introduce you to today. Uh, so the first one um, we're going to call table of contents and it's somewhat like a tree structure uh, but with also different types of links between them. You'll see how it works.

当你面对一个「长得都差不多」的大 schema,里面有一堆非常相似的表,你怎么搞清楚该怎么正确地 join 它们?所以今天我们会给你们介绍三种具体的「形状」。第一种,我们叫它「目录(table of contents)」,它有点像一个树形结构,但节点之间还带有不同类型的连接。你们等下就会看到它是怎么运作的。


[6:35]

Um this is going to be used on the unstructured data. Uh on the unstructured data site as well, we have something called themes. Um and what themes does is it surfers surfaces global patterns inside of your data that might not be apparent um or you might not know about beforehand. So uncovering unknown patterns and groupings. Um, and then the third one which when we take the course we're actually going to go over first is this connection shape

这个是用在非结构化数据上的。在非结构化这一侧,我们还有一个叫「主题(themes)」的东西。themes 的作用是浮现出你数据里那些不那么显而易见、或者你事先根本不知道的全局模式,也就是发掘未知的模式和聚类。然后第三个——虽然我们上课时实际上会先讲它——是「连接(connection)」形状,


[7:02]

which is essentially a semantic layer on top of your data warehouse. And so how many people just a show of hands here are familiar with like graph and graph rag. Okay so a fair number of you then how many people in here have used neo forj before?

它本质上是架在你数据仓库之上的一个语义层。那么,举个手看看,这里有多少人熟悉 graph 和 GraphRAG?好,还挺不少的。那这里有多少人以前用过 Neo4j?


[7:25]

Okay. Also like uh maybe 30 40% something like that. So this room seems like you know fair fairly uh fairly familiar with graph basics but I'll I'll go over it just here really quick. So Neoraj we're we're a graph intelligence platform. We have a graph database at at at the heart of it. Right. Um and basically when we talk about what a graph is we're talking about it from a property graph data model perspective.

好,也差不多有 30%、40% 这样。所以这个房间对 graph 的基础应该算相当熟悉了,不过我还是快速过一下。Neo4j 我们是一个 graph 智能平台,核心是一个 graph database(图数据库),对吧。基本上,当我们说「什么是 graph」的时候,我们是从属性图(property graph)数据模型的角度来讲的。


[7:53]

And so everything inside of our database and our analytics platform is modeled to these three different types of elements which is nodes that represent people, places and things. Relationships which are the verbs or associations between those things. So like person owns car, person drives car,

所以在我们的数据库和分析平台里,所有东西都被建模成三类元素:node(节点),代表人、地点和事物;relationship(关系),也就是这些事物之间的动词或关联,比如「人拥有车」「人驾驶车」、


[8:12]

person is related to other person. And then we have properties that go on these both the relationships or sometimes we'll call those edges and the nodes which are sometimes called vertices. And those can be anything from strings, they can be numbers, they can be dates, they can be vectors,

「某人和另一个人有关系」。然后是 property(属性),它可以挂在 relationship(有时候我们也叫 edge,边)上,也可以挂在 node(有时候叫 vertex,顶点)上。属性可以是字符串、数字、日期,也可以是 vector(向量),


[8:32]

um um all sorts of things. And so the basic idea with a graph, right, is as you start adding data to it, it's almost like a bunch of pre-joined tables and everything's already interconnected and you can hop between all the nodes very easily. Um, and that provides a ton of benefits that we'll see we'll see later in the course.

各种各样都行。所以 graph 的基本思路是:当你不断往里加数据,它几乎就像是一堆预先 join 好的表,所有东西天然就互相连着,你可以非常轻松地在各个 node 之间跳转。这带来了大量好处,我们在课程后面会看到。


[8:52]

So, with that in mind, why don't you go ahead and get started and go over into these if um if you haven't heard before and I'll actually just show you here really quick. Um inside of the workshop lakehouse, have it open here. I should probably restart this guy really quick.

那么带着这些背景,如果你之前没听说过,不妨现在就开始动手,进到这些页面里去。我这就快速给你们演示一下。呃,在 workshop lakehouse 里面,我这边已经打开了。我可能得先快速重启一下这个家伙。


[9:16]

Start my code base. Um, but basically it should take you to this page. Um, this shouldn't wouldn't say continue course for you if you if it's your first time going to it. It would say enroll to take the course. So, you have to go through that path to enroll to take the course.

启动我的 code base。基本上它应该会带你到这个页面。呃,如果你是第一次访问,它不会显示「继续课程」,而是会显示「注册以参加课程」。所以你得走那条路径,先注册报名这门课。


[9:34]

Um, and then basically the whole thing from here on out is just going to be me guiding you through this course and kind of showing you all the different uh things that we can do. Um, there's about 17 lessons in here. We've broken it up by the different shapes. Um, but basically if you go into the course,

然后从这里开始,接下来基本上就是我带着你们走完这门课,给你们展示我们能做的各种事情。呃,这里面大概有 17 节课。我们是按不同的「形状」来拆分的。基本上,如果你进到课程里,


[9:54]

you can kind of click through here. I go over the scenario a little bit. But while I'll do that, I think what you should do is jump to this environment section. So basically, you go down to the bottom here. go to your environment and um from here you can say open code space and in fact I will open a new code space here just so that I can I can walk through it uh live with you as well.

你可以在这里一路点下去。我会稍微讲一下场景。不过我讲的同时,我觉得你们应该先跳到「环境」这一节。基本上你往下拉到最底下,进到你的 environment,然后从这里点「open code space」。事实上我这边也会新开一个 code space,这样我就能跟你们一起现场走一遍。


[10:27]

Yeah, of course. Um, go back to here. Why don't you all go ahead and capture that? Yes. So it'll create a code space in your own GitHub if you want. There's a link to the repository. You could run it locally. Um just keep in mind that we've set it up so that it will auto start in the code space. So there's another um bash script that you can run to run it locally and it will it will set everything up for you. Um but it it

当然可以。呃,回到这里。你们都先把这个拍下来吧。对。它会在你自己的 GitHub 里创建一个 code space。这里有一个仓库链接,你也可以在本地跑。呃,只是要注意,我们的配置是让它在 code space 里自动启动的。所以本地跑的话还有另一个 bash 脚本可以执行,它会帮你把所有东西都配好。不过这


[11:17]

depends a little bit on your local environment. You have to make sure you have cloud code installed um and all these sorts of things. So, we'd recommend that you use a Code Space one if if you're comfortable doing it locally. I I I've run it locally all the time before, too.

多少取决于你本地的环境。你得确保装了 Claude Code 之类的东西。所以我们建议你用 Code Space;当然如果你觉得本地更顺手也行,我自己以前也一直在本地跑。


[11:38]

All right, everyone's everyone's copied the uh QR code. I'll give it another 10 seconds or so and then I'll I'll move back over. Yep. Ben and Ryan can help you out if if you don't have this one. Okay. So, when you start up as well, there's some workshop credentials here. Uh I'll I'll be uh blasting these keys uh a after the course. I'll leave them up for just a little bit. Um but basically,

好,大家都把二维码拍好了吧。我再给 10 秒左右,然后我就切回去了。嗯。如果你还没拿到这个,Ben 和 Ryan 可以帮你。好。另外,启动的时候这里有一些 workshop 凭证。呃,课程结束后我会把这些 key 作废掉。我先让它留一会儿。基本上,


[12:17]

give it some time to start up. Uh- once it once it does, what you'll do is you'll go into your environment file here. And maybe I can make this just a little bit bigger so it's easier to see. You will place your anthropic key in here as well as the uh the BigQuery key.

给它一点时间启动。一旦起来了,你要做的就是进到这里的环境变量文件。我把字体调大一点,这样好看清。你要把你的 Anthropic key 填进去,还有那个 BigQuery key。


[12:38]

The other thing that you're going to want to place in uh if you go back to the course um is down here, it should give you a um Neo Forj credential. So, this is accessing our graph database. They've been all pre-provisioned, so it's going to be different for each one of you. You're going to have a different one. So, copy the one from your screen.

另外还有一个你需要填进去的东西——如果你回到课程页面,往下这里,应该会给你一个 Neo4j 的凭证。这是用来访问我们的 graph database 的。它们都已经预先分配好了,所以每个人的都不一样,你会拿到你自己那份。所以请复制你自己屏幕上的那一份。


[13:06]

Um, and you're going to put it inside of this environment file as well. Um, and it it will take a while, like you see with mine, the environment file hasn't uh quite popped up yet. So, it can take a few minutes for that all to go through. Um, so we'll we'll give that a little bit of time to start up.

呃,你也要把它放进这个环境变量文件里。呃,这会花点时间——就像你看到我这边,环境变量文件还没冒出来呢。所以整个过程可能要几分钟。呃,那我们就给它一点时间启动。


[13:32]

Hopefully, it won't take too long here. While that's running, I'll I'll give a little bit of an overview here of what we're going to do in the next step. So, basically, once we get this set up and and we get Claude up and running, um we're going to walk through um the first shape for our day, which is this connection semantic layer shape. And then after that we'll follow it with the table of contents one that we were

希望别等太久。趁它在跑,我先大致讲一下下一步我们要做什么。基本上,等我们把环境配好、把 Claude 跑起来之后,我们会先走今天的第一个形状,也就是「连接」语义层这个形状。之后我们再接着讲前面提到的


[14:02]

talking about and the communities one. So we'll be working with the structured data first and then moving into the into the unstructured bits. Um once for those of you who do have it loaded basically you have your terminal down here. You can go ahead and call Claude. Um, we'll be working through Claude Code as our agent.

「目录」形状,以及社区(communities)那个。所以我们会先处理结构化数据,然后再进入非结构化的部分。呃,对于那些已经加载好的同学,基本上你的终端就在下面这里,你可以直接调用 Claude。呃,我们会用 Claude Code 作为我们的 agent。


[14:30]

Um, the internet is a little bit slow too, which is expected as you go through with this as well. I should um Oh, here we go. We'll set up show you what this looks like here in a second. Anthropic key as well as BigQuery one. All right, perfect.

呃,网络有点慢,这在这种场合也是意料之中的。我应该……哦,好了,出来了。我们马上给你们看看这长什么样。Anthropic key,还有 BigQuery 那个。好,完美。


[15:46]

And if it does this where it's trying to get you to do a subscription, basically what you do is you'll go plus, open up a new terminal window, and then it should here give you the option to select the enthropic key. Just go through and press enter. Yes.

另外,如果它跳出来让你去开订阅,你要做的基本上就是点「+」新开一个终端窗口,然后它应该就会给你选项,让你选择用 Anthropic key。一路走下去按回车就行。好。


[16:05]

And then you want to use the MCP server here. So you're going to select yes. And I'm going to skip the tour. And then I can ask that question just to make sure it can connect to Neo4j. And for me, I already have some nodes in because I'm working with um a graph that's been cached already for my previous course.

然后这里你要用到 MCP server,所以选 yes。我这里跳过引导教程。接着我就可以问那个问题,确认一下能不能连上 Neo4j。对我来说,里面已经有一些 node 了,因为我用的是之前课程里已经缓存好的一个 graph。


[16:51]

Um for you, it will it will come up as an empty sandbox. So, one thing that uh I want to note here too in terms of how we're going to be working with Neo forj and with cipher, which is our graph query language. I think we've come to a point now where a lot of us aren't h hand typing uh our own code line by line anymore. We're using agents to help us build things obviously. And so in this workshop it's going to be the same thing. Basically

对你们来说,打开会是一个空的 sandbox。这里我还想说明一点,关于我们接下来怎么使用 Neo4j 和 Cypher——Cypher 就是我们的 graph 查询语言。我觉得我们现在已经到了这样一个阶段:很多人都不再一行一行手敲代码了,大家显然都在用 agent 来帮忙构建东西。所以这个 workshop 里也是一样。基本上


[17:24]

we're going to be using something called the Neo Forj CLY um which is a CLY tool that will allow your agent to run uh queries directly against the database. When I was working inside of Claude, you saw it running there for a second. And that also ships with a Neoraj cipher skill. Um, that allows the agent to sort of leverage how to put queries together in the most modern way, as well as something called a GDS skill, which

我们会用到一个叫 Neo4j CLI 的东西,这是一个命令行工具,能让你的 agent 直接对数据库执行查询。刚才我在 Claude 里操作的时候,你们看到它跑了一下。它还自带一个 Neo4j Cypher skill,能让 agent 知道该怎么用最现代的方式把查询组织起来。另外还有一个叫 GDS 的 skill,


[17:52]

stands for graph data science. And we'll use that for some of the algorithms when we get to the theme section. Um, in terms of just helping us do agentic coding. So, I'm not going to have you guys handw write really any um cipher query code. I'm going to show you throughout this workshop how to use an agent to help you write that with with specs that are created. Um and hopefully that'll be useful for you more on on the

GDS 代表 graph data science(图数据科学)。等我们讲到 theme 那一节的时候,会用它来跑一些算法。这些都是为了帮我们做 agentic coding。所以我基本上不会让大家手写任何 Cypher 查询代码。整个 workshop 我会演示怎么用 agent、配合写好的 spec 来帮你生成这些代码。希望这对你们日常工作也有帮助,


[18:18]

day-to-day if you ever decide to use uh cipher um whether it be with Neo forj or not. Um, so how is everybody doing now? I I'll take a second here to stop. Go ahead and raise your hand if you have any questions here. Or are we all good? I have one question over there.

如果你们以后要用 Cypher 的话——不管是不是配合 Neo4j 用。那么大家现在进展怎么样?我停一下。有问题的话请举手,或者大家都没问题?那边有一个问题。


[18:55]

Yeah. So the question is, can I use cursor? And so the answer to that is you should be able to. Yes. I haven't personally tested it, but if you go up to the top, there should be a link to the repository that we're using. Um if not, when you go to open code, yeah,

好的。问题是:能不能用 Cursor?答案是应该可以。我个人没有实际测试过,但如果你往上翻,应该能看到我们用的那个仓库的链接。如果没有的话,等你打开 open code 的时候,对,


[19:15]

you can clone the repository here. So if you follow that link, you could clone it. You can get it locally. Um and then you'll see here there's directions. There's a shell file that you can run that will set everything up. You can take a look at what's in there. It's it's a bunch of basic stuff. Okay.

你可以在这里克隆这个仓库。所以只要顺着那个链接,你就能 clone 下来,在本地跑起来。然后你会看到这里有说明文档,有一个 shell 脚本,运行它就能把环境全部配好。你可以看看里面的内容,都是些很基础的东西。好。


[19:31]

Awesome. Yeah. And just a little bit about the structure of this thing. I forgot to mention um there's a lot of stuff inside of the uh inside of uh the the workspace here. Um obviously inside of Cloud we have our skills file. Um, so I've gone ahead and pre-written a skill here. Um, that covers basically a lot of what we're going to do. Uh, inside of here as well,

太好了。再简单说一下这个东西的结构,我刚才忘了提——这个 workspace 里面有很多内容。显然在 Claude 里面我们有 skills 文件。我已经预先写好了一个 skill,基本涵盖了我们要做的大部分事情。这里面还有,


[20:02]

there's uh a file called uh outline and search and theme. So, these are going to be the shapes that we'll be working with. um to query BigQuery there's just a very small um well that this is actually a a database query file but the run SQL one this one for querying bigquery just a very simple um um basically way of uh reaching out to BigQuery and just calling it with a uh with um a simple SQL command. Um and inside of these files as we'll see later

有几个文件叫 outline、search 和 theme。这些就是我们要用到的几种「形状」(shape)。至于查询 BigQuery,这里有一个很小的……嗯,其实这个是数据库查询文件,但是那个 run SQL 的文件,就是用来查 BigQuery 的,非常简单,基本上就是连到 BigQuery,然后用一条简单的 SQL 命令去调用它。这些文件里面,我们后面会看到,


[20:35]

there's there's places to fill in a spec. So that's going to be where we do some of our agentic coding. Um, but everything in here, if you wanted to like load this into BigQuery yourself, there's something for that. Um, if you want to go ahead and run this on data bricks, there's directions here to like load it into into data bricks and use Genie and AI search there if you want to compare. Um, and then this is just a sim

都留有填写 spec 的地方。那就是我们做 agentic coding 的部分。这里面什么都有,比如你想自己把数据加载到 BigQuery 里,有对应的脚本;如果你想在 Databricks 上跑,这里也有说明教你怎么把数据加载进 Databricks,用 Genie 和 AI search 来做对比。然后这个只是一个符号


[21:00]

link over back to the uh cloud files. Um, there's solution scripts as well. Also if you get stuck for example where for some reason the coding agent can't you know create the right query um you you can copy stuff from there. Um and then all of the source data too for creating both the structured and unstructured data is in here um including our uh PDF sources. So if you go in here we have our bulletins and our manuals and everything. Um this is

链接(sym link),指回到 Claude 的那些文件。这里还有解决方案脚本。比如你卡住了,coding agent 因为某些原因生成不出正确的查询,你就可以从那里复制。另外,用来生成结构化和非结构化数据的所有源数据也都在这里,包括我们的 PDF 源文件。所以进到这里,能看到我们的公告(bulletin)、手册(manual)等等。这个是


[21:31]

showing you the uh the raw PDF. But if I go to the corpus, which is the the markdown version, uh you'd be able to to actual re actually read the contents inside of them. Um so anyway, with all of that, we'll go ahead and hop along to our first section here, which is going to be connections.

展示原始的 PDF。但如果我去看 corpus,也就是 markdown 版本,你就能真正读到里面的内容了。好,说完这些,我们就往下走,进入第一节:connections(连接)。


[21:57]

So, we have basically our data inside of BigQuery right now. And in fact, if I go back to my overview, just to show you what it looks like really quick. This is the uh schema. Uh so, we've kept it very simple for this course. Obviously, in the real world, you'd have a much bigger schema than this. Um but we have our work orders, uh which is sort of like the middle or the star of the scheme.

现在我们的数据基本上都在 BigQuery 里。我回到 overview 页面,快速给大家看一下它长什么样。这是 schema。为了这门课程我们把它做得很简单,显然现实世界中你的 schema 会比这个大得多。我们有 work orders(工单),它算是整个 schema 的中心,或者说星型结构的中心。


[22:23]

you have your vehicles, um the different uh DTC codes, procedures, uh work order parts, and then the parts themselves. Um and they join together in a very simple primary foreign key pattern. Um we'll get into our documents later in the course. It's just basically a a big bucket of of PDF files.

还有 vehicles(车辆)、各种 DTC 故障码、procedures(维修流程)、work order parts(工单零件),以及 parts(零件)本身。它们之间通过很简单的主键-外键模式关联起来。文档部分我们课程后面再讲,那基本上就是一大堆 PDF 文件。


[22:44]

And so what we want to do here is to help our agent be able to effectively join everything together, we're going to give it a graph representation of a semantic layer. And to build that semantic layer, we're going to use something called Neoarta. Um Neoarta is a labs project. So at Neo Forj, we have our core engineering, which is like our,

所以我们在这里想做的是:为了帮助 agent 能有效地把所有东西关联起来,我们要给它一个语义层(semantic layer)的 graph 表示。而要构建这个语义层,我们会用到一个叫 NeoCarta 的东西。NeoCarta 是一个 labs 项目。在 Neo4j,我们有核心工程团队,负责比如


[23:10]

you know, cloud SAS platform, um our graph analytics stuff and all of that. But we also have labs projects where we move a little bit faster. Um, and inside of that labs project, we created this neocarta. Um, and what it does is it creates a metadata graph. Um, also although we're not leveraging it here,

云端 SaaS 平台、graph analytics 之类的东西。但我们也有 labs 项目,在那里我们迭代速度会快一些。在这个 labs 项目里,我们做出了 NeoCarta。它的作用是创建一个元数据 graph(metadata graph)。另外,虽然我们这次不会用到,


[23:27]

you can add business terminology and business processes and a whole bunch of other things into a graph structure. Um, and if you think about it, people, you know, these words ontology and semantic layer and all these things are kind of they're thrown around a lot lately,

你还可以把业务术语、业务流程以及其他一大堆东西加进这个 graph 结构里。你想想看,ontology(本体)、semantic layer(语义层)这些词最近被到处乱用,


[23:42]

especially this summer. And so the the basic uh mental paradigm that I use to kind of think through these is an ontology helps something interpret and reason about the data. A semantic layer helps something understand what the consistent agreed upon terms are so it can query the data accurately. Um and then at Neo Forj we have other things like virtual graph. There's a link in here in the course which will actually give you a sort of a a federated graph

尤其是今年夏天。我自己用来理清这些概念的基本思维框架是这样的:ontology 帮助某个东西去解释和推理数据;而 semantic layer 帮助某个东西理解大家一致认可的统一术语是什么,这样它才能准确地查询数据。另外在 Neo4j 我们还有别的东西,比如 virtual graph。课程里有个链接,它能给你一个联邦式(federated)的 graph


[24:07]

view of your SQL data that you can um that you can query over directly with cipher. We won't be using that here, but there there is um sometimes uh value in in giving sort of a graph view of the schema. Um but anyway, what Neo Carter will do is it we'll see in this next section um uh it has an MCP server basically that will sort of suck all the data in from the um from BigQuery just the metadata and it will create this graph.

视图来看你的 SQL 数据,你可以直接用 Cypher 在上面查询。这次我们不会用到它,但有时候把 schema 以 graph 视图呈现出来确实是有价值的。总之,NeoCarta 做的事情,我们下一节就会看到——它本质上有一个 MCP server,会把 BigQuery 里的数据、其实只是元数据,全部抽取进来,然后生成这个 graph。


[24:39]

Um and that's what you do in the next section. So, I'll go ahead and jump to there. And I've already run this script. But basically, when you see this script, what you're going to do is you're going to copy it. And then you're going to go over to here. And then, um, you're going to bring up your terminal, which our terminal.

这就是下一节你们要做的事。我直接跳到那里。我已经跑过这个脚本了。基本流程是:你看到这个脚本以后,把它复制下来,然后切到这边,把终端调出来,就是我们的 terminal。


[25:02]

And then, uh, I like to just open a new window. That's not my cloud window. Um, and you're going to allow paste. And then you're going to go ahead and run that. I've I've already run it. Um so I won't do it again, although it should be item potent. Um but basically once you run that, uh it'll create this um these six tables with these five reference join paths. And once you've done that,

然后我习惯新开一个窗口,不用我的 Claude 那个窗口。接着允许粘贴,然后运行它。我已经跑过了,所以就不再跑一遍了,虽然它应该是幂等(idempotent)的。基本上你跑完之后,它会创建出这六张表,以及这五条参考 join 路径。做完这一步之后,


[25:29]

you can come back in here and even just run this to check. And it's kind of cool what it looks like. Um this is this is what it looks like inside of the graph database. And basically what you'll see is you have this uh this node in the middle that represents your um your well first your database and then your schema and then if you follow that out it'll be has table and then you'll have your table and then you have your columns. Um

你可以回到这里,运行这个来检查一下。它的样子还挺酷的。这就是它在 graph 数据库里的样子。基本上你会看到中间有一个 node 代表你的……嗯,首先是你的 database,然后是 schema,顺着往外走会是 HAS_TABLE 这条边,然后是你的 table,再往下是 columns。


[25:58]

and those can also optionally have these representative values on top of them. Um, and so these join paths, basically when the agent reads that, um, if I was to go down here and try, so I have, um, a little example prompt down here and I can go into my window and go back to Claude and I can ask a question like this.

而且这些还可以可选地带上代表性的取值(representative values)。所以有了这些 join 路径,agent 读到之后……比如我往下翻试一下,我这里有一个示例提示词,我可以切到我的窗口回到 Claude,然后问这样一个问题。


[26:28]

Which vehicle received part IC 2042? And I in this case I'm going to tell it to use the the MCP server and the warehouse schema and all this to to go and grab everything then use the Python script. Um inside of the skill itself it it has directions for this as well. Um,

哪辆车装了零件 IC 2042?这种情况下我会告诉它使用 MCP server 和 warehouse schema 等等,让它去把所有信息抓过来,然后用那个 Python 脚本。skill 本身里面其实也有相应的说明。


[26:45]

but what you'll see it'll do of this. It'll call that MCP server for Neo Carta and it's it's going to read that metadata semantic layer graph that we have in Neo forj. And so if you notice the thing that we're doing here really is we're not using the graph to copy the data over. There's not like an ETL into graph. What we're doing is we're using the graph as a semantic layer. So we're just pulling metadata about the columns

你会看到它这么做:它会调用 NeoCarta 的那个 MCP server,去读取我们存在 Neo4j 里的那个元数据语义层 graph。注意我们这里真正在做的事情——我们并不是用 graph 把数据复制过来,这里没有往 graph 做 ETL。我们做的是把 graph 当作一个语义层。所以我们只是把列的元数据拉过来,


[27:20]

and the rows and all of these things and we're going to use that to inform a text to SQL query once it uh decides to go ahead and run here. How is claude working for everyone right now? Are we seeing a lot of uh slow clouds? Fast. Yeah, slow.

还有行的信息等等,然后用这些去指导一次 text-to-SQL 查询,等它决定要执行的时候就会跑起来。大家现在的 Claude 跑得怎么样?是不是很多人的 Claude 都很慢?有快的吗?嗯,慢。


[27:51]

Okay. Interesting. I'll give it a little bit of time here. I don't know why it would be go moving. Huh. A little bit slow.

好吧,有意思。我这里再等它一会儿。我不知道为什么会这样。嗯,是有点慢。


[28:31]

Yeah. Does everyone have Yes.

对。大家都有……有的。


[28:48]

Yeah, sure. So, the question is, can I recount how I got to here while we're we're waiting for Claude to come along? So, basically, um the QR codes that are being passed around, there's two here. uh one of them is uh to enroll in the course. So the course that I'm going through is online. It's freely available so anyone can go uh once once you show up inside of the course basically you'll see a button here to enroll. Um so you

好,当然可以。问题是:能不能趁等 Claude 的功夫,重新讲一遍我是怎么走到这一步的?基本上,现在传下去的那些二维码,这里有两个。其中一个是用来注册课程的。我正在讲的这门课程是在线的,免费开放,任何人都能上。你进到课程页面以后,会看到这里有一个 enroll(注册)按钮。所以你


[29:17]

have to sign in and enroll in the course. Uh once you do that you'll you'll click through and you'll make it to this environment section. Um, and when you do that, uh, it'll take about five minutes or so for everything to to populate and come up. Um, if you're using code spaces, you have the option to run it locally. Um, but then once you do that, essentially, um, you're going to come in here and you're going to get

需要先登录,然后注册这门课。做完之后,你一路点下去就会来到 environment(环境)这一节。到这一步之后,大概要等五分钟左右,所有东西才会准备好并加载出来。如果你用的是 Codespaces,你也可以选择在本地运行。做完这些之后,你基本上就进到这里,然后去拿


[29:43]

your credentials and you're going to grab the anthropic API key, um, as well as the, uh, the BigQuery key here. Um, and then once you do that, you'll be able to um access all the things in the workshop. Basically, um, you're using, as you saw here, um, oh, cool. We're actually making progress. Um, we're using cloud code inside of inside of code spaces here.

你的凭证,取到 Anthropic API key,以及这里的 BigQuery key。拿到之后,你就能访问这个 workshop 里的所有东西了。基本上,就像你们刚才看到的……哦,不错,真的有进展了。我们是在 Codespaces 里面用 Claude Code。


[30:15]

Sorry, can you repeat that really quick?

抱歉,能再说一遍吗?


[30:21]

Yes. The Neo Forj sandbox is um if I if I were to go back to that environment section um you see the button here to open code space and then you should have down here um your own sandbox credentials. So, it won't be these exact ones because it's it's different for every participant, but you'll go ahead and copy those. And then once you're in here, uh you'll be able to um you have your environment file. And inside of

对。Neo4j sandbox 的话,如果我回到 environment 那一节,你会看到这里有个按钮可以打开 Codespace,然后下面这里应该有你自己的 sandbox 凭证。所以不会和我这些一模一样,因为每个参与者的都不一样。你把它们复制下来。然后等你进到这里之后,你会有一个环境变量文件(environment file)。在


[30:58]

your environment file where you put your anthropic and your BigQuery key, you'll also put the Neo Forj ones. If you don't see a environment file and you see nothing or you see the end.ample, example, it just takes it a while to populate that. There's other things that are running. It's running another shell script to kind of create everything for you and um populate all the different files.

这个环境变量文件里,除了填 Anthropic 和 BigQuery 的 key,你还要把 Neo4j 的那几个也填进去。如果你没看到环境变量文件,什么都没有,或者只看到 .env.example 这个示例文件,那只是因为它需要一点时间才能生成出来。后台还有别的东西在跑,它在运行另一个 shell 脚本,帮你把所有东西创建好,把各种文件都填充好。


[31:22]

So, my my hunch, and someone can scream at me if I'm wrong, is that if Anthropics this slow right now, it's more of a anthropic problem and not necessarily a one person's account problem. Um, if it continues being bad, I might try to provide another key from a different account, but I don't actually think that that's what the problem is. I think it's it's

我的直觉是——如果我说错了大家可以纠正我——如果 Anthropic 现在这么慢,那更可能是 Anthropic 那边的问题,不一定是某个人账号的问题。如果一直这么糟糕,我可能会从另一个账号再提供一个 key,但我其实不觉得问题出在这里。我觉得是……


[31:45]

Yeah. Well, yeah, but this should also be running remotely, right?

嗯。不过话说回来,这个应该也是在远程跑的,对吧?


[31:50]

Is it Is it just the Wi-Fi? Okay.

是不是只是 Wi-Fi 的问题?好吧。


[31:55]

Yeah. Well, the funny thing is I would think in Well, never mind. I don't know. It's running in Codespaces, but maybe Code Spaces is still using the the local uh the local Wi-Fi. Um, okay. So, so we're back here. Basically, what basically what it did to grab this list is um it it went through and it it basically read um that metadata graph um and then it created the um the different uh queries that it wanted to run. You

对。有意思的是,我本来以为……算了,我也说不好。它是在 Codespaces 里跑的,但可能 Codespaces 用的还是本地的 Wi-Fi。嗯,好,我们回到这里。它要拿到这个列表,基本上做的事情是:它去读取了那个 metadata graph,然后生成了它想要执行的那些查询。你


[32:28]

can see it ran that the small SQL uh Python file um and it it did all of its selects and it joins to be able to to pull together uh these tables that list out um the the answer to that question essentially. Um if I was to go back the question here was uh what would we ask?

可以看到它运行了那个小的 SQL Python 文件,做了一系列 select 和 join,把这些表拼到一起,本质上就得出了那个问题的答案。如果我回过头看,这里的问题是——我们当时问的是什么来着?


[32:50]

Go back to the connection shape. Um yeah, which vehicles received, you know, these these separate parts over here. Um and and it it provided that list um down here by by make apparently um and then vehicle count. Um so that's the basics of this first shape. And you can imagine that as your data starts to grow and you start to get more and more tables, it's very useful to have this semantic layer um to then help essentially guide how everything

回到 connection shape。对,就是哪些车辆装配了这边这些特定的零部件。它在下面给出了这个列表,看起来是按品牌(make)分的,然后是车辆数量。这就是第一种 shape 的基本情况。你可以想象,当你的数据越来越多、表越来越多的时候,有这么一个语义层(semantic layer)来指导所有东西怎么


[33:26]

joins together. Uh in this case, it used a primary foreign key uh relationship inside of the information schema inside of BigQuery uh to enforce that. Neoarta can also use things like query logs or you can sort of manually put together how you want things to join with other terminology and metrics.

join 在一起,是非常有用的。在这个例子里,它用的是 BigQuery information schema 里的主外键(primary/foreign key)关系来做约束。NeoCarta 也可以用查询日志(query logs)之类的东西,或者你也可以手动去定义你希望这些东西怎么 join,以及配上其他术语和指标。


[33:45]

Um, are there any uh questions about this first sort of connections graph semantic layer shape that we're working with?

关于我们正在讲的第一种 shape,也就是 connections graph 语义层,大家有什么问题吗?


[33:57]

Yes, I have one over here. Neo for you. So the question was can you give us more context into the Neo Forj near quarter like what we're doing strategically um so what Neo Forj is doing strategically and this right here so Neo forj as a company right we're for a long time you've been able to import data into our database and represent things as a graph which is very useful if you're trying to run these graph specific queries right so if you have a

好,这边有一位。Neo4j 相关的问题。问题是:能不能多讲讲 Neo4j 和 NeoCarta 的背景,比如我们战略上在做什么。那我说说 Neo4j 战略上在做什么,以及这块东西。Neo4j 作为一家公司,长期以来你都可以把数据导入我们的数据库,把它表示成 graph——如果你要跑那些 graph 特有的查询,这非常有用。比如说你有一条


[34:40]

supply chain for example and you need to find an optimal route from point A to B and you have to do these like variable length, you know, shortest path type of calculations. A graph database is very good for that. But the other thing that a graph database is really great at doing is not just making those complicated queries run faster, but actually providing a view or a representation of the data that allows an agent to understand how tables might

供应链,你需要找到从 A 点到 B 点的最优路径,需要做那种可变长度的最短路径计算,graph 数据库在这方面非常擅长。但 graph 数据库真正厉害的另一点,不只是让这些复杂查询跑得更快,而是它能提供一种数据的视图或者说表示方式,让 agent 能理解表和表之间可能


[35:06]

interrelate. So that even though the end join it might only be a three or four hop join, you might have had to understand hundreds of tables or something to be able to arrive at that conclusion. And so we're focused a lot right now on these concepts of ontologies and semantic layers and what we're calling virtual graph where the focus isn't only just etling data in but about okay maybe you want to keep your data where it is and you still want to

如何关联。所以哪怕最终的 join 可能只有三四跳,你可能得先理解上百张表才能得出那个结论。所以我们现在很大精力放在 ontology(本体)、semantic layer 这些概念上,以及我们叫做 virtual graph 的东西——重点不只是把数据 ETL 进来,而是:好,也许你想让数据待在原地,你还想继续


[35:32]

use SQL and and access it in in the ways you have been but you need some way to sort of guide the agent to be able to do that correctly. Um so Neo Carta is a labs project. We also uh just recently um in preview have released uh something called virtual graph which will be similar but it essentially gives you um almost like these push down cipher queries. So it gives you a graph schema of your of your database and then it

用 SQL、按你原来的方式去访问它,但你需要某种方式来引导 agent 正确地做这件事。所以 NeoCarta 是一个 Labs 项目。我们最近还发布了一个预览版的东西叫 virtual graph,它跟这个类似,但本质上它给你的几乎是下推式(push down)的 Cypher 查询。它给你一个数据库的 graph schema,然后


[36:00]

allows you to run cipher directly um which can be useful because then you sort of have the view down at the query interface level and not just at the metadata semantic layer level. So, there was a table in there um in the previous section that was kind of explaining that, but hopefully that's helpful so you understand kind of where where we're going with things.

允许你直接跑 Cypher。这很有用,因为这样你在查询接口层就有了这个视图,而不只是在 metadata 语义层上有。前面那一节里有张表就是在解释这个的,希望这些能帮你们理解我们的方向。


[36:23]

Yeah. Yep. Grap and Yes. And we're doing a ton of stuff inside of context graphs as well and and and um and all of those layers. Yep. All right. Awesome. Um any other questions? Okay, I have one question over here.

对,没错。Graph 还有——对。我们在 context graph 以及所有这些层面上也做了大量工作。对。好的,很棒。还有其他问题吗?好,这边有一个问题。


[37:03]

Yeah. So, so the question is um why would we not take our OOLTP data which is right now is this warehouse of vehicle orders and and stuff and just migr or or some subset of it. Why wouldn't we just push that into the graph? Well, I think in a lot of cases that's easier said than done, right? Because if you think about a lot of production use cases, you can have terabytes of data which gets updated continuously.

对。问题是:我们为什么不把 OLTP 数据——现在这些车辆订单之类的数据仓库——直接迁移,或者迁一部分进 graph 里?我觉得很多情况下这说起来容易做起来难。因为你想想很多生产环境的场景,你可能有 TB 级的数据,而且还在持续更新。


[37:32]

And so, right, if you had to put that into the graph, you would then have to find some way of syncing that while you move everything in. Um, and there's might be a lot of extra properties or things that you might need to build a custom ETL where there's some things that you might not want in the graph.

所以如果你要把这些放进 graph,你就得在搬迁的同时想办法做同步。而且可能有很多额外的属性,或者有些东西你得写自定义的 ETL,因为有些内容你并不想放进 graph 里。


[37:49]

And then a lot of times what people end up not really understanding until they get in those situations is the security posture. Because if you have sensitive data that's in one system, it might not even though you could physically move it in, you you might not actually be able to for security reasons take that data and just move it into another database.

还有一点,很多人往往要等到真的碰上了才意识到,就是安全合规的问题。因为如果你有敏感数据在某个系统里,哪怕你物理上能把它搬过去,出于安全原因你实际上可能并不被允许把那份数据直接搬到另一个数据库里。


[38:08]

Um so so there's various reasons why you might not necessarily want to move your data over. Uh where you do get an advantage of the ETL over just this metadata and semantic layer graph is if you have graph queries that need that performance, right? So like that supply chain case that I had earlier like if you're doing these really large recursive joins and you need them to happen very quickly or you need to run graph algorithms maybe you need to

所以有各种各样的原因让你未必想把数据搬过去。那么相比只做 metadata 加语义层 graph,ETL 的优势在哪儿呢?就是当你有 graph 查询、确实需要那种性能的时候。比如我前面说的供应链那个例子,如果你要做超大规模的递归 join,并且需要它很快出结果,或者你需要跑图算法,可能你需要


[38:34]

produce graph embeddings or you need to do some sort of clustering or something in that case it makes a lot more sense to to bring the data in. Um and then as we'll see in the next sections when your data is already unstructured uh there's benefits to bringing it into a graph because then you can give it a graph structure.

生成 graph embedding,或者要做某种聚类之类的,那种情况下把数据搬进来就非常合理了。另外,我们在接下来几节会看到,当你的数据本来就是非结构化的,把它放进 graph 是有好处的,因为这样你可以给它赋予一个 graph 结构。


[38:51]

All right. I saw one other question over here. Yes.

好的。我看到这边还有一个问题。请讲。


[39:03]

So the the question is does it rebuild the graph every time you initialize Neo Carta? And the answer to that question is no. What it will do is basically when uh we we ran that build connection script um it went ahead and ran an MCP server function that ports the data in um and then once it's ingested into Neoarta just the metadata then that stays there. So like for me I didn't rerun that build connections right because I just took the course last

问题是:每次初始化 NeoCarta 的时候,它会不会重建整个 graph?答案是不会。它做的事情基本上是:当我们运行那个 build connection 脚本时,它调用了一个 MCP server 的函数,把数据导进去;一旦这些 metadata 被 ingest 进 NeoCarta,它就一直在那儿了。比如我自己,我就没有重跑 build connections,因为我昨晚


[39:33]

night. Um so like for me I and I probably could have because I I think it's an item pot and load but I just didn't want to deal with it. So, it was there for me and then I just ran the queries on top of it. Yep. All right. Any any other questions?

刚过了一遍这个课程。我其实也可以重跑,因为我记得这个加载是幂等(idempotent)的,只是我懒得折腾。所以对我来说数据已经在那儿了,我直接在上面跑查询就行。好,还有其他问题吗?


[39:50]

I have one over here at the end. Yes. Yes. You

后排这边有一位。请讲。对,你说。


[40:04]

the the Neo Forj cipher skill. Um, so I don't know if I might be able to uh grab it really quick. Um, I can show you Neo forj cipher skill. But this is these are all developed internally by our team. So if you go here actually through the Neo Forj CL,

关于 Neo4j 的 Cypher skill。嗯,我看看能不能快速调出来。我可以给你们看看 Neo4j Cypher skill。不过这些都是我们团队内部开发的。你如果通过 Neo4j CLI 进去,


[40:33]

you can download a ton of these skills. So if you see here, there's there's the cipher one, but then we also have the different ones for like agent memory, for cloud infrastructure, for different drivers, like if you're using Java or Go or or all those things. Um, so I can't necessarily say exactly how we landed on this particular skill MD, but what I will tell you is it was created by us and our internal teams who are like

就可以下载一大堆这样的 skill。你看这里,有 Cypher 那个,我们还有各种其他的,比如 agent memory 的、云基础设施的、各种 driver 的——比如你用 Java 或者 Go 之类的。所以我没法准确说出我们是怎么最终定出这个 SKILL.md 的,但我可以告诉你的是,它是我们和内部团队做的,他们


[40:59]

always up to date on the latest, you know, cipher 2526. Um, so so it it it has all of the latest. So I' I'd encourage you to use the Neo Forj CLI because it will basically you can load one or multiple of these skills and you'll get the latest knowledge there because it otherwise what what it's going to do is it's going to go on the internet and get the Stack Overflow questions from like four or five or six years ago and it's

一直紧跟最新的东西,比如 Cypher 25、26。所以它包含了所有最新的内容。我建议你们用 Neo4j CLI,因为你可以加载一个或多个这样的 skill,拿到的是最新的知识。否则的话,它会跑到互联网上去抓四五六年前的 Stack Overflow 问答,然后


[41:26]

going to give you outdated bad cipher basically.

给你一堆过时的、糟糕的 Cypher。


[41:31]

CLI

CLI


[41:33]

it is. Yes, it is loaded in. And uh if you look at the um the install file that we use inside of dev container, uh you'll see um where we installed the Neo Forj agent client, the skill. So the commands in there um that that basically allowed you to do that.

是的。对,已经加载进去了。如果你看我们在 dev container 里用的那个安装文件,你会看到我们安装 Neo4j agent client 和那个 skill 的地方。里面的那些命令基本上就是让你能做到这一点的。


[41:54]

All righty, I'll take one more question and then I think I'll have to move on to the next section unless there isn't another one. Yes, right there. um bronze so like the different like delta tables where um so I mean we're we're exploring basically um the traditional like SQL structured data warehouse and then NeoAR Carta is really customer-driven for us. It was made by our field team. So as customers come in and they have questions about less un

好,我再回答一个问题,然后如果没有别的,我就得进到下一节了。好,就是那位。嗯,bronze——比如说不同的 Delta 表之类的。我们现在探索的基本上是传统的、SQL 结构化的数据仓库,而 NeoCarta 对我们来说是非常客户驱动的,它是我们的 field team 做出来的。所以随着客户进来,问到关于更


[42:42]

less structured data um and even things like documents have been considered for Neo Carta and so there's this backlog that's forming right now. Everything's moving very fast. I think I mean Ryan is is somewhere in the room and he might be able to to speak on sort of what's what's the next thing for Neo Carta, but I know we're working right now on a data bricks connector um and maybe some metric views and other things.

非结构化数据的问题,甚至像文档这类东西也在 NeoCarta 的考虑范围内,所以现在正在形成一个 backlog。一切都在飞快推进。我想——Ryan 应该就在这个房间里,他也许能讲讲 NeoCarta 下一步会做什么,但我知道我们现在在做一个 Databricks connector,可能还有一些 metric view 之类的东西。


[43:16]

All right, awesome. So meet with uh with Ryan afterwards and then everyone watching on YouTube I guess can meet with Ryan afterwards but by then we'll we'll have we'll have an amazing Neil Carta. So all righty. So so let me go ahead and move on for sake of time because we're coming 15 minutes on the halfway point and it might be nice to take a few minute break at some point in the middle just to just to give ourselves a little bit of a breather. So

好,很棒。那大家会后可以找 Ryan 聊;至于在 YouTube 上看的各位,我猜你们也可以会后找 Ryan,不过到那时候我们的 NeoCarta 应该已经很厉害了。好,那么为了赶时间我先继续往下讲,因为我们离过半点还有 15 分钟,中途找个时间休息几分钟应该不错,让大家喘口气。那么


[43:41]

now I'm going to jump to the next section. And so that was working with structured data and now I want to work with unstructured data. And this is going to be the documents part. So we have if we go over to our little uh our little uh source data stuff here. So in this case I'm going to read them locally to load them. Um but you know it's reading from cloud storage reading locally. Um there there often isn't a huge difference. Um basically we have

现在我要跳到下一节。刚才那部分是处理结构化数据,现在我想处理非结构化数据,也就是文档这一块。我们过去看一下我们的源数据。在这个例子里,我打算从本地读取来加载它们。不过你知道,从云存储读和从本地读,通常差别不大。基本上我们有


[44:16]

these bulletins in PDFs. Um we have manuals which are considerably longer. Um and then we have recalls as well. And because it's giving me this funky view, I also have them in uh in this raw corpus where they're they're inside of um these uh markdown files here. Um, but basically, uh, there are these multi-section documents. Um, and you'll see here that they have, uh, this linking that they do to go over to different sections in other documents.

这些 PDF 格式的技术公告(bulletin),还有手册(manual),那些要长得多,另外还有召回通知(recall)。因为这个视图显示得有点怪,我这边也把它们放在了一个 raw corpus 里,就是这些 markdown 文件。基本上,这些是多章节的文档。你会看到它们里面有链接,会跳转到其他文档的不同章节。


[44:50]

So, they're all kind of interlin together as they refer to one another. Um, so, for example, right inside of inside of this manual, um, it talks about these different platform codes for this ABS system, um, and diagnostic troubleshooting, right? and it will link over to there. All of this data, by the way, is I simulated. It's not sensitive or or anything like that. Um, but it's meant to simulate kind of a real world

所以它们相互引用、彼此串联在一起。比如说,在这份手册里,它讲到这个 ABS 系统的各种平台代码,还有诊断排障,然后它会链接过去。顺便说一下,所有这些数据都是我模拟生成的,不是敏感数据什么的。但它是为了模拟一个真实世界的


[45:15]

uh system where you have these manuals and bulletins and recalls of all these different vehicles. And um what we want to accomplish with this basically is we want to give our agent a table of contents. And and by a table of contents, I mean something that looks like this thing. So, what I want to allow the agent to do here is not just search for key terms. We'll give it search too, but I don't know if has anyone heard of like page index? Anyone

系统,里面有各种车型的手册、技术公告和召回通知。我们基本上想达成的目标是:给我们的 agent 一份目录(table of contents)。我说的目录,就是长得像这个东西。所以我希望让 agent 能做的不只是搜关键词——我们也会给它搜索能力,但不知道大家有没有听说过 page index?有人


[45:45]

know who they are? Maybe some of you. So, there's so there's this idea of navigation through your documents where basically almost like a human if you look at the table of contents inside of a book and not I'll call this outline. I'll call it table of contents. I need to get the terminology a little bit better aligned. But the idea here is that you can sort of read the indentations and you can see how you have a document or in this case a

知道他们是谁吗?可能有些人知道。所以这里有一个「在文档中导航」的想法,基本上就像人一样,如果你去看一本书的目录——我这里叫它 outline,也叫它 table of contents,术语上我需要再统一一下。但这里的想法是,你可以读懂那些缩进,你能看出你有一个文档,或者在这个例子里是一个


[46:13]

library. You have the bulletins which is a subfolder. Then you have these documents underneath. And so you have this containment tree. And then in addition to that containment tree that gives you all these sections, you also have these links that'll take you to a different document. So it's more than just a table of contents that's a tree that goes down. It also has these links,

library(资料库)。你有 bulletins,这是一个子文件夹。然后下面挂着这些文档。所以你就得到了这样一棵包含关系树(containment tree)。除了这棵给你划分出所有 section 的包含树之外,你还有这些 link,能把你带到另一份文档去。所以它不只是一个自上而下的树状目录,它还带有这些 link,


[46:37]

right? And these links and basically everything in here have a URI. And the idea is that if we can come up with a good graph representation, we can basically make it so that these URIs which are hierarchical. So you see like you have your technical library here,

对吧?而这些 link,以及这里面基本上所有东西,都有一个 URI。想法是,如果我们能设计出一个好的 graph 表示,我们基本上就能让这些 URI 是分层的。你看,比如你这里有 technical library,


[46:55]

right? And then slashbulletins. So, like if I'm down in, you know, this TSB recall notice, I know that it's part of bulletins, um, or safety bulletin, rather, not a recall, and it's part of a technical library, just like it's also linking to a manual here. Um, and you can see the name of the manual, then it's inside of the manual subfolder. Um,

对吧?然后是 /bulletins。所以,如果我在下面这个 TSB 召回通知这里,我就知道它属于 bulletins——呃,应该说是 safety bulletin,不是召回——而且它属于 technical library,就像它同时也链接到这里的一份手册一样。你能看到手册的名字,它就在 manual 这个子文件夹里面。


[47:19]

and if I have that, I can take that and I can plug it in and I can actually grab a node from the graph that has like the raw text or or even a link to the raw text. And so, and then if I uh plug this in and we'll see how this works with outline later, I can actually get these sub trees so I can like dig in and drill down on these different pieces of content. Um, that's the general idea with this. And so what this gives the

有了这个之后,我就可以把它拿过来、填进去,然后真的从 graph 里抓出一个 node,这个 node 上有原始文本,或者甚至只是一个指向原始文本的链接。然后如果我把这个填进去——我们待会儿会看到它和 outline 是怎么配合的——我就能拿到这些子树,这样我就能往下钻取、深入到这些不同的内容片段里。这大体上就是这套东西的思路。所以这给 agent 带来的能力是,


[47:44]

agent to do is not just search like vector search or or lexical search but actually kind of traverse through the documents in a sense. Um and the data model that we're going to use in this case we're actually going to import the data into graph. Um I have a I have a quick question for everyone. So, everyone here is knows Graphrag. When they think of Graphrag,

不只是做 vector search 或者词法检索那种搜索,而是在某种意义上真正地在文档之间穿行。我们这里要用的数据模型,我们实际上会把数据 ingest 进 graph 里。我先快速问大家一个问题。在座的都知道 GraphRAG。当你们想到 GraphRAG 的时候,


[48:12]

and maybe someone can raise your hand and I'll I'll pick on someone for a second. I want you to tell me what you think like the pro the the building the graph process is for Graph Frag. So, I don't know if anyone wants to wants to volunteer if they're familiar. You're going to raise your hand.

也许有人可以举个手,我点一位来说说。我想让你告诉我,你觉得 GraphRAG 里「构建 graph」这个过程是什么样的。不知道有没有人愿意来讲讲,如果你熟悉的话。你要举手了。


[48:30]

Yes. Yeah. And they build they build the they build the links from from the So there's different ways, right? And and you said I think in in your when you talk about creating um a graph, you build the links from the data that's inside of the documents, right?

对。是的。他们会去构建那些 link,从……所以是有不同做法的,对吧?我想你刚才在讲创建一个 graph 的时候说的是,你是从文档内部的数据里去构建这些 link 的,对吧?


[49:02]

And so graph rag oftentimes we have an entity extraction piece to it as well. So we can use an LLM to kind of say hey like extract the different parts the different vehicles from the from the graph and have them be separate entities and then have those links to like the original document chunks and all all of this stuff. And that's great and that works really well. Um here what I'm going to show you is something that's

所以 GraphRAG 里我们通常还会有一个 entity extraction 的环节。我们可以用 LLM 来说,嘿,把这些不同的部分、不同的车型从文档里抽出来,让它们成为独立的 entity,然后把它们链接回原始的文档 chunk,等等这一整套。这很好,而且效果也确实不错。但我这里要给你们展示的,是一种


[49:31]

much more lightweight and I think maybe to the point that you were making um we're actually going to use a deterministic loading. So if you see here the way that this graph structure is going to work is you're going to have your library and then you're going to have this containment tree that reflects what we saw. And all this is really doing is it's breaking down, you see the folders, the document, and then these documents have sections and they can

轻量得多的做法,而且我觉得这可能正好呼应你刚才提的那个点——我们实际上要用一种确定性的加载方式。所以你看这里,这个 graph 结构的工作方式是:你有你的 library,然后你有这棵反映我们刚才所见结构的包含树。它做的事情其实就是层层拆解,你能看到文件夹、文档,然后这些文档有 section,而且它们可以


[49:59]

have multiple sections that will also nest underneath each other at different section levels. Um, and then you'll have um next section links. So you get a concept of ordering as well as linked to which uses in this case we're using um these named links uh sort of like you might see inside of an Obsidian vault a little bit um but you can also do it with hyperlinking and other things depending on depending on what your data

有多个 section,这些 section 还会以不同的层级互相嵌套。然后你还会有 next section 这样的 link。所以你既有了顺序的概念,也有了 linked to 这种关系——在这个例子里我们用的是这些具名 link,有点像你在 Obsidian 仓库里可能会看到的那种。不过你也可以用超链接或者别的方式来做,取决于你的数据


[50:27]

looks like. So, it's all deterministic with the containment tree for one and then the ordering of the sections as well as the links. Um, and the benefits of having a deterministic load like this is number one, it's going to be item potent. Um, so like you're not relying on an LLM in the beginning. Um, it's often going to be a little bit faster.

长什么样。所以这一切都是确定性的:一是包含树,然后是 section 的排序,再加上这些 link。这种确定性加载的好处,第一,它是 idempotent(幂等)的。你在一开始就不依赖 LLM。而且它通常会快一些。


[50:48]

Um and so if you already have documents which have a lot of inherent structure to them and a lot of interlinking um sometimes just to get something up and running quickly it can be very beneficial to approach a graph like this where I'd say it's more of a lexical or document structure graph rather than like a full like entity uh you know extraction type of pipeline to create a graph.

所以,如果你已经有了本身结构性很强、互相之间链接也很多的文档,有时候为了快速把东西跑起来,用这种方式来做 graph 是非常划算的。我会说这更像是一个词法层面的、或者说文档结构的 graph,而不是那种完整的、靠 entity extraction 的 pipeline 来构建 graph。


[51:11]

Um and it it explains here what I just said with the different link types. Um and then and and like I said before these URLs carry a hierarchy. So every node for example that that head node at the top here for the library will get its URI manuals will have that plus back slashmanuals um going down sort of the containment tree to the to the file name and then the different sections and um this is the script you're going

这里也解释了我刚才说的那些不同的 link 类型。然后就像我之前说的,这些 URL 是带层级的。所以每个 node,比如最上面这个代表 library 的头节点,会拿到它的 URI;manuals 就是在它基础上再加 /manuals,顺着包含树一路往下走到文件名,再到不同的 section。然后这就是你们要


[51:44]

to copy to load it in. So basically again you'll take this, you'll go over to here. Um you'll open your bash terminal. Now I believe I've already done this so I'm not going to run it again even though I I'm pretty sure it's item potent. Um but basically once you run that um what you'll end up getting is and I'll just make sure it's actually in the graph. It is for me. You'll end up getting that uh that model that we just

复制过去执行加载的脚本。所以还是一样,你把这个拿走,切到这边来,打开你的 bash 终端。我记得我已经跑过了,所以我就不再跑一遍了,虽然我挺确定它是 idempotent 的。但基本上你一旦跑完这个,你最后会得到——我先确认一下它确实已经在 graph 里了,对我来说是有的——你最后会得到我们刚才看到的


[52:10]

saw. It should only take Well now it might take a while if the Wi-Fi is slow. although I don't think it should. Um, but it'll create this graph for you where you have again your top of the folder structure which is our technical library, our folders, and then you can see we have our documents followed up here into our different sections. And if I were to zoom in here, um, you'll see we have the has links in the next section, but then we

那个模型。它应该只需要……嗯,现在如果 Wi-Fi 慢的话可能要花点时间,不过我觉得应该不至于。它会给你建出这个 graph:你还是有文件夹结构的顶层,也就是我们的 technical library,然后是我们的文件夹,接着你能看到我们的文档,再往上到不同的 section。如果我在这里放大,你会看到我们有 has links 和 next section,但我们


[52:37]

also have the links too. And I can go in and click on these and you can see kind of how the thing grows and everything kind of links to each other. And now that we have that, um, basically in the next section, we can go ahead and build our tool. Um, basically the script that goes along with our our custom autofix skill that we've created uh to um to actually like create that outline shape that we saw in the beginning, the

同时也有 links to。我可以点进这些看看,你能感受到这东西是怎么长出来的,所有东西怎么互相链接。有了这个之后,基本上在下一节里我们就可以去构建我们的工具了。也就是配合我们做的那个自定义 autofix skill 的脚本,用来真正生成我们一开始看到的那个 outline 形状,


[53:10]

the table of contents to hand our agent. And so if I go ahead to that next section, um, and again, like I said, what I want to do here is I'm going to go ahead and copy this. This is the this is the prompt. Um, where if I go ahead and feed that to my claude,

也就是要交给我们 agent 的那个目录。所以我切到下一节,还是像我说的,我这里要做的是把这段复制过来。这就是那个 prompt。如果我把它喂给我的 Claude,


[53:35]

it should go ahead and get to work. And because I know that's going to take a while, I'm going to have it start running here. But basically, what it's doing as it goes through um is if I go under and I look at my skills, um you'll see that I left inside of here um this wired query. This is this is the query that it needs to create um to be able to create that outline shape. And I've instructed it inside of the language here to use the Neo Forj

它就会开始干活了。因为我知道这要跑一会儿,我就先让它在这儿跑起来。它跑的过程中做的事情基本上是——如果我进去看我的 skills,你会看到我在这里面留了这个连线好的 query。这就是它需要写出来的那个 query,好让它能生成那个 outline 形状。我在这里的说明文字里指示它去用 Neo4j 的


[54:17]

cippher skill and reference a spec in doc outline format.m MD. So this is generally how I like to write more complicated cipher queries, more complicated graph queries is I'll write a spec for it. Um, and so if I go to my docs folder, you'll see I'll have my specs in here and I have my outline format spec where I say this is the shape, right? And then it it talks here about how it needs to be human readable and uh yes, we will allow you to edit

Cypher skill,并参考 doc outline format.md 这个 spec。所以这大体上就是我写更复杂的 Cypher 查询、更复杂的 graph 查询时的习惯做法:我会先给它写一份 spec。那么如果我打开我的 docs 文件夹,你会看到我的 spec 都在这儿,我有这个 outline format 的 spec,我在里面说明「这就是那个形状」,对吧?然后它这里还讲到它需要是人类可读的,以及——是的,我们会允许你编辑


[54:51]

and um and some of the optional arguments it needs to have. So, right, because oftent times it's like pseudo code that we'll have and we don't know exactly how to put it together. Um, um, in this case, I did give it the hint about having a variable length path query uh to basically put everything uh put everything in order.

还有它需要支持的一些可选参数。对吧,因为很多时候我们手上是一堆伪代码,我们并不确切知道该怎么把它们拼起来。在这个例子里,我确实给了它一个提示,让它用变长路径查询(variable length path query),基本上就是为了把所有东西都按顺序排好。


[55:20]

[snorts] And then it will say when it's done. And so basically what you'll get out of that is if I go back here and it will talk a little bit about like one of the keys to this you'll see in the file is it does this thing here which it goes has which is that containment edge and you see this star 0.25 it's actually parameterized inside of the inside of the query. So like uh if I go back to my um me go to where it changed

[吸鼻子]然后它做完了会告诉你。基本上你从中得到的结果是——如果我回到这边,它会讲一点,你在文件里会看到,这里有一个关键点,就是它用了 has 这个东西,也就是那条包含关系的 edge,你看到这个 *0..25,它其实在 query 里是参数化的。所以,如果我回到我的……让我切到它改动的地方


[56:06]

um if I go back to here it's it's parameterized inside of this string. So it gives you like an option to how deep you want to traverse. Um and that that's sort of the graphy nature of kind of putting together this table of contents. There's there's a second query in here too which is much simpler which just grabs the links after that. Um so that basically uh once that's done I can go ahead and um copy I'll run the whole thing. If if you just

如果我回到这里,它在这个字符串里是参数化的。所以它给了你一个选项,可以决定你想遍历多深。这在某种程度上就是拼出这份目录时那种「图的味道」。这里面还有第二个 query,简单得多,就是在那之后把这些 link 抓出来。所以基本上,等它跑完,我就可以复制过来,我把整个东西跑一遍。如果你


[56:36]

run the script without any depth parameter or without a specific URI, it's going to return everything in the graph. So it wouldn't be the way that I'd I'd run it with with a larger graph here. We end up with because we only have a couple hundred documents. Um so I can get away with it here. Um, but if I ran that,

不带任何 depth 参数、也不指定具体 URI 就跑这个脚本,它会把 graph 里的所有东西都返回回来。所以如果 graph 更大,我不会这么跑。我们这里最后能这么干,是因为我们只有几百份文档,所以我在这儿可以偷这个懒。但如果我跑那个的话,


[57:00]

um, it'll it'll bring back a ton. And, but you can kind of see here, um, zoom out. So, you can see, right, that it gave me this. And you can see how a lot of these documents again, it's not just the structure, it's like the links between all of them that it's that it's providing.

它会带回来一大堆东西。不过你在这里大致能看出来——我缩小一点——你能看到,对吧,它给了我这个。你能看到这么多文档,还是那句话,它给出的不只是结构,更是它们彼此之间的那些 link。


[57:19]

Um and then what you can do with this right is you can give it an optional depth parameter. So for example, if I just said depth equals one. Uh go you don't like clearing to the bottom and I say depth one. Then if I did that um it'll go ahead and just go one down. And then the other cool thing with this too is you can provide it a specific URI.

然后你能拿它做的事情是,你可以给它一个可选的 depth 参数。比如说,我就写 depth 等于 1。呃,你不喜欢一路拉到底部——我写 depth 1。那如果我这么做,它就只会往下走一层。另外一个很酷的点是,你还可以给它一个具体的 URI。


[57:51]

So if I saw you know a specific URI here that I wanted to give it um then it will basically start at that Falcon 2.0 no document just like you know if I um if I was here now and say okay well actually you know I want to do um I want to do it but I want to do it for something that I see up above like maybe this links to you know this um coil identification thing I can go ahead and snag that URI and then I can put it in the script

所以如果我看到这里某个我想传给它的具体 URI,那它基本上就会从那个 Falcon 2.0 文档开始。就像,如果我现在在这儿,然后说,好,其实我想做的是——我想做,但我想针对上面看到的某个东西来做,比如这个链接到的这个线圈识别(coil identification)的东西——我就可以把那个 URI 抓过来,然后放进脚本里,


[58:29]

and then it will basically give me you know that document and everything it links to. So you can say see how an agent can go through now and sort of grab these URIs and kind of traverse all through the through the graph and give itself these hierarchical and and linked views.

然后它基本上就会给我那份文档以及它链接到的所有东西。所以你能看到,一个 agent 现在可以怎样去抓取这些 URI,在整个 graph 里到处穿行,并给自己构造出这些分层的、带链接的视图。


[58:50]

Um the okay yeah why don't we why don't we take a second for questions. So we have one in the middle here. Um well so later in the course I'll in the workshop here for some of the last questions I do instruct the agent to say exactly what steps it followed. Um are you seeing it use something different?

好,我们要不要花点时间答疑。中间这位有一个问题。这样,在课程后面——在这个 workshop 后面的一些问题里——我确实指示了 agent 明确说出它到底走了哪些步骤。你是看到它用了别的方式吗?


[59:34]

Right.

对。


[59:37]

Yeah. So it should be apparent in the tool call history. Um so you can you can look at the different tool calls that it made because in the case for example of connections it's calling an MCP server. So those should be inside of your tool calls. And the question by the way for everyone if if you couldn't hear um it was just how do you know that the agent's actually calling the things that it's supposed to right yeah so so for

是的。所以这在工具调用记录里应该是能看出来的。你可以去看它发起的那些不同的 tool call,因为比如说在 connections 这个例子里,它调用的是一个 MCP server。所以那些调用应该都在你的 tool call 里能看到。顺便把问题复述一下给大家,如果你们刚才没听清:问题就是,你怎么知道 agent 真的调用了它该调用的那些东西,对吧?是的,所以对于


[1:00:02]

cloud code for example you can look at the tool call history um and informally what I'll do in this course is for some of the prompts I'll I'll just direct it to say like hey like clearly like say like what steps and logic you use to be able to to answer this question. Um, you know, obviously for a production system,

比如 Claude Code,你可以去看工具调用记录。而在这门课里我非正式的做法是,对某些 prompt,我会直接指示它说:嘿,明确地讲清楚你用了哪些步骤和什么逻辑来回答这个问题。当然,对于生产系统来说,


[1:00:22]

you would you would want to look at the logs and everything, but for here and for learning, I I that's how I'll I'll expose it. Um, is there are there any other questions about uh trees outlines?

你会想去看日志之类的东西,但在这里、为了教学,我就是这么把它暴露出来的。关于树、outline,还有别的问题吗?


[1:00:37]

Yes. One right here.

好,这边这位。


[1:00:45]

Yeah.

是的。


[1:00:47]

Yeah. So, um it is uh all the code for that should be in uh oh yeah, we'll need to restart. Um oh, that was a different codebase. Thank Thank goodness. All right. Um yes, inside of load, if you go to load documents, this is this is just the cipher queries.

对。所有相关的代码都应该在……哦对,我们得重启一下。呃,那是另一个代码库,谢天谢地。好的。是的,在 load 里面,如果你打开 load documents,这里面就只是一些 Cypher 查询。


[1:01:16]

And then you have your parse corpus um down here. So yeah, it's basically just using regx in here to try to find, you know, the the right stuff. In this case, the documents are already fairly well structured, right? So mileage may vary depending on what your document sources are. In this case, right, if you have a lot of manuals that are often structured the same way, then you can get away with sort of regex and plain

然后下面这里是 parse corpus。基本上就是用正则表达式来找出我们想要的那些内容。在这个例子里,文档本身已经结构化得相当好了,对吧?所以具体效果会因你的文档来源不同而不同。像这种情况,如果你有大量结构基本一致的手册,那你用正则加一些简单的


[1:01:43]

NLP techniques. If it's messier, you might have to use like um you know, gler or different types of looms and there's stages of complexity for that. All right. Yes, we have a question over here. Yes. Yep. Can be done with any model. The only thing that Neoarta and any of these tools well neo I'll start with neocarta neocarta is really responsible for creating that semantic layer metadata graph and providing you an MCP endpoint

NLP 技术就够用了。如果文档更乱一些,你可能就得用 GLiNER 或者别的 LLM,这里面有不同的复杂度层级。好,这边有个问题。是的。对,任何模型都可以。NeoCarta 和这些工具唯一……嗯,我先说 NeoCarta,NeoCarta 真正负责的是创建那个语义层的 metadata graph,并且给你提供一个 MCP endpoint。


[1:02:24]

right so as long as you can access that MCP server it's model agnostic okay so that was that was a question about basically we're using opus 4.8 here and can we use any other type of model and the answer to that question is is yes we're using a skills framework and then in and then for Neo Carta we have MCP um Neo Carta also has a CLI so there's there's other ways to access it too um I have one question all the way in back there

所以只要你能访问那个 MCP server,它就是模型无关的。好,刚才那个问题是说,我们这里用的是 Opus 4.5,那能不能换成其他模型,答案是可以的。我们用的是 skills 框架,然后 NeoCarta 这边我们有 MCP。NeoCarta 还有一个 CLI,所以还有其他访问方式。后面那位有个问题。


[1:03:09]

So the question is around if if I'm understanding correctly, how do you decide what to name your relationships and how descriptive they have to be and who is naming the relationships. So in this case, we have a very deterministic load. So we're deciding ahead of time what to name our relationships. Um, and that is a bit of a a taste or a judgment call, right? Uh, I tried to keep it simple here where has is very simple.

这个问题是,如果我理解得没错的话,你怎么决定 relationship 该叫什么名字、名字要多有描述性,以及是谁在给 relationship 命名。在我们这个例子里,load 过程是完全确定性的,所以我们是提前就决定好了 relationship 叫什么。这多少有点靠品味或者主观判断,对吧?我这里尽量保持简单,比如 HAS 就非常简单。


[1:03:40]

It's a containment relationship. I also know that I'm going to want containment going multiple levels. So, I probably want a common name for all of those because if I have like has folder, has document section, it makes the cipher complicated. Um, so that's sort of why I kept that naming very short. Same with links too. Um the more granular you get with naming um basically the more sort of detailed uh your agent can be and

它是一个包含关系。我也知道我会需要多层级的包含关系,所以我大概想给所有这些层级用一个统一的名字,因为如果我搞成 HAS_FOLDER、HAS_DOCUMENT_SECTION 这样,Cypher 就会变得很复杂。所以这就是我把命名保持得很短的原因。LINKS 也是一样。命名粒度越细,你的 agent 和


[1:04:08]

your end tools can be in putting a query together. But then the more complicated your data model becomes. And so it's a little bit of a thing that you have to manage because if you end up in like a production scenario where you have hundreds of different types of relationships um that might get hard to manage inside of a context window and it can lead to models maybe not having the easiest time putting those cipher queries together with something like the

下游工具在拼查询的时候就能越精确。但与此同时,你的数据模型也会变得越复杂。所以这是个需要你自己权衡的事情,因为如果你到了生产场景,有几百种不同类型的 relationship,那在 context window 里可能就很难管理了,也会导致模型不太容易顺利地把那些 Cypher 查询拼出来,比如用


[1:04:32]

Neo Forj CL or humans, right? When we're trying to, you know, write our own uh sort of parameterized queries. Yep. Yeah, that's true too. Um, and we see that right that that's the whole semantic layer argument too with on the warehouse side is you can have these metrics views and other things which use business terminology that might be different from sort of the the physical data model that you have, right? Um, and

Neo4j 的 CLI,或者人来写的时候也一样,对吧?我们自己写那些带参数的查询的时候。对。那个说法也没错。我们也看到,这其实就是数仓那边所谓语义层的论点——你可以有这些 metrics views 之类的东西,它们用的是业务术语,可能跟你实际的物理数据模型不一样,对吧?


[1:05:02]

so that's that's part of the reason why we have these metadata graphs to to help manage that sort of different or that delta between them. I'll take a couple more questions and then we'll we can move on. Yep.

所以这也是我们为什么要有这些 metadata graph 的原因之一,就是帮你管理这两者之间的差异、这个 delta。我再回答几个问题,然后我们就继续往下讲。请讲。


[1:05:27]

Yeah, that's an interesting question. Um, I think it would be naive to call it a replacement because most of the people that I see using this will eventually incorporate some sort of hybrid vector retrieval or full text search work I do. I use at least full text search um, with this sort of stuff. So, I don't think it's a replacement, but as we'll see later in the course when we get to some of the estate level questions. Um,

这个问题很有意思。我觉得把它称为“替代”是有点天真的,因为我见到的大多数用这套东西的人,最终都会加入某种混合的 vector 检索或者全文搜索。我自己做的工作里,至少会用全文搜索配合这套东西。所以我不认为它是替代品,但我们后面课程讲到那些资产级别(estate level)的问题时会看到,


[1:05:54]

there's certain things that semantic search is not as great at answering that really require document navigation to understand how like everything connects together, right? Um, so big examples would be like proving something doesn't exist. Like how do you do that with semantic search, right? or understanding um even like how you know there might be different parts related to a specific document or something. You know, you

有些问题语义搜索并不擅长回答,那些问题真的需要文档导航,需要理解各个部分是怎么连接起来的,对吧?最典型的例子就是证明某个东西不存在。你用语义搜索怎么做到这一点?或者理解某个具体文档可能关联到哪些不同的零部件之类的。你


[1:06:22]

could in theory do like a bunch of vector queries to keep pulling passages, but there's more, you know, chance for error that you might not catch the right document and all this sort of stuff. Um so yeah, not the question was if it's a replacement for semantic search and I don't think it is. No.

理论上可以做一堆 vector 查询不断去捞段落,但出错的概率更大,你可能捞不到正确的文档,诸如此类。所以是的,刚才的问题是它是不是语义搜索的替代品,我认为不是。


[1:06:41]

Um, one more in uh back there. So, um the question is do we redefine the types of nodes um as well as the relationships and so is this a question of like how do you decide what the node labels should be? Um, yeah. And again, it's it's again, it's a data modeling question, right? Like there could be a situation where we're looking at the data model that we just went over. Um,

后面那位再来一个问题。这个问题是,我们是不是也会像定义 relationship 一样,去预先定义 node 的类型?所以这是在问怎么决定 node label 应该是什么?嗯,对。这同样是个数据建模的问题,对吧?比如说,我们刚才过的那个数据模型,可能会有这样一种情况……


[1:07:20]

well, it's not it's not here. It's in my uh it's in my last one. um when I look at the uh when I look at the shape um zoom out of that a little bit where it might be useful to have like an actual parts node or something right um here for this type of use case I'm letting the agent kind of infer parts and extract them from the documents and then use that to go find something from the metadata graph and query SQL or the or

嗯,不在这里,在我上一个里面。当我看那个 shape 的时候……稍微缩小一点……在某些场景下,可能确实有必要真的建一个 parts 节点之类的。但在我们这个用例里,我让 agent 自己去推断零部件、从文档里把它们抽取出来,然后拿这个去 metadata graph 里查东西、去查 SQL,或者


[1:07:53]

the other way around. Um, and that works for me. So, you know, it's still as as things evolved, like two years ago, I probably would have said you would need that like additional node. Now, it seems like agents are smart enough where maybe you don't all the time anymore. Um and but for some use cases, you know,

反过来也行。这对我来说够用了。所以你看,这个东西一直在演进,两年前我大概会说你必须要有那个额外的 node。现在看起来 agent 已经足够聪明,可能不是所有时候都需要了。但对某些用例来说,


[1:08:13]

especially like we see this in life sciences for example, where there's these really specific ontologies around, you know, different types of molecules and medications and things and you you really do need that represented in the graph or else it just your inference isn't going to make sense. So I hate to say it depends on use case. I don't want to be that person, but it it sort of does in a sense. All right. Um,

比如我们在生命科学领域就经常看到,那里有非常具体的本体(ontology),涉及各种分子、药物之类的,你就真的需要把这些在 graph 里表示出来,否则你的推理根本讲不通。所以我不太想说“这要看用例”,我不想当那种人,但某种意义上确实就是这样。好的。


[1:08:38]

oh, okay. Yeah, let me let me move on only because we're a little bit over halfway through and I have some more course to to get through here. So, um, let's go to So, we went over all of this stuff. Um, oh yes. So, we need to build our uh search query. So, similar to what we just did and to the the question that the gentleman had over here, uh we want to be able to um we want to be able to build um some sort of search as well to help augment this.

哦,好。我还是继续往下讲吧,因为我们时间已经过半了,后面还有不少内容要讲。那么,这些我们都过完了。哦对,我们还需要构建搜索查询。所以跟我们刚才做的类似,也呼应这边这位先生提的问题,我们希望能够构建某种搜索能力来作为补充增强。


[1:09:18]

So, um this is going to be much more simple, this search uh file. Um basically, I'm just going to use a lucine style index. Uh, and I I'm not going to use vector search here. It's the Neo Forj itself has um a lucine full text search index,

这个搜索文件会简单很多。基本上我就用一个 Lucene 风格的 index。这里我不打算用 vector 搜索。Neo4j 本身就带有 Lucene 全文搜索 index,


[1:09:38]

but and I'm going to do and if you see here, it's basically doing the same thing where I have the uh this this wire that it has to fill in. Um, and it's going to use a cipher skill to do that. While it's it's running there, you're allowed to edit. Um,

但是我要做的是……你看这里,基本上是同样的套路,我有这个 WHERE 需要它去填。它会用一个 Cypher skill 来完成这件事。它在跑的时候,你是可以编辑的。


[1:09:58]

we might as well wait till it's done. All right, it's going to finish up. So, basically, uh, if you see here, it's the way this works is you call the full text index. Um, it's indexed on the document and the section nodes. Uh, so every node that has text in it, like we don't put folders through through this lucine index. Um, and then it will give you a score. Um and then because we have everything in those URI format um we can

我们不如等它跑完。好,它快跑完了。基本上,你看这里,它的工作方式是你调用全文 index。这个 index 建在 document 和 section 节点上。所以每个含有文本的节点——像 folder 我们就不放进这个 Lucene index。然后它会给你一个分数。又因为我们所有东西都是 URI 格式的,我们还


[1:10:28]

also scope it by subtree um basically with like a starts with filter. Uh so if I was to go back into this uh file um and I looked at the way that this query run I have this wear statement. Um so I'm applying this as a post filter basically where I'm doing the full text search first. Um and then I am filtering it by the URI. So I can say like only search under like this section of the manuals basically. Um so it just helps

可以按子树来限定范围,基本上就是用一个 starts with 过滤。所以如果我回到这个文件里,看这个查询是怎么跑的,我有这条 WHERE 语句。我把它当作后置过滤器来用,也就是先做全文搜索,然后再用 URI 来过滤。这样我就可以说,只在手册的这一部分下面搜索。所以这就帮我


[1:10:57]

me refine my my search a little bit with those hierarchical URIs. Uh, and the other thing that I do, um, and and I do this because I work like in in the the research role that I have, I work with so many different AI models and on so many different platforms. Like you wouldn't think about it, but like choosing like and wiring like options to be like work with like six different vendors for vectors is kind of tough. So

用这些层级化的 URI 稍微收敛了一下搜索范围。另外我还会做一件事,我这么做是因为在我的研究岗位上,我要跟非常多不同的 AI 模型、在非常多不同的平台上打交道。你可能想不到,但是要把选项接通、让它能同时兼容六家不同的 vector 厂商,是挺麻烦的。所以


[1:11:23]

what I often do for this is I'll use something called semantic expansion. Um and basically what it does is it it instructs the AI model to say hey like use your world knowledge if someone asks for you know an engine shuttering that it might be a misfire or something else too. So it and because this is a lucine index um I can do that. So for example right I can say like misfire or rough idle and it will search for either of

我经常做的是用一种叫做语义扩展(semantic expansion)的东西。它的作用基本上是指示 AI 模型说:嘿,用你的世界知识——如果有人问“发动机抖动”,那可能指的是 misfire 或者别的什么。因为这是个 Lucene index,我可以这么做。比如我可以搜 misfire 或者 rough idle,它会把两者


[1:11:51]

those. Um, if I go back to my bash script and I run that now that the uh thing's filled it in and then you'll see it'll give me my uh documents with with the scoring for for the different ones. And it I can also search as I was saying before under recall. So I can search for coils under just the recall uh subfolder library. So that would be the idea here.

都搜出来。如果我回到 bash 脚本,现在它已经把内容填好了,我跑一下,你会看到它会给我返回文档以及各自的评分。而且像我刚才说的,我也可以限定在 recall 下面搜索。所以我可以只在 recall 这个子文件夹库下面搜 coils。大概就是这个思路。


[1:12:18]

Um, so you can see it's still leveraging uh that hierarchical containment shape, the tree, um, but it's doing it through the URI um, ID structure so that it doesn't have to do as much graph reversal. It just sort of filters down to to things underneath.

所以你能看到,它仍然在利用那个层级化的包含结构、那棵树,但它是通过 URI 这个 ID 结构来实现的,这样就不用做那么多图遍历,直接就过滤到下面的那些东西。


[1:12:37]

Um, all righty. Awesome. Um, are there any quick questions just around the search piece there? The lucine search. We have one question over here. Yeah. So the question is we just loaded the documents into the into lucine and then we're searching underneath different sections appropriate sections. Yeah. So basically if it might make more sense if we actually took a look at the the query that we're running here. So here's the I'll make it this big.

好的,很好。关于搜索这一块,也就是 Lucene 搜索,有没有什么简短的问题?这边有一个问题。对。问题是说,我们只是把文档灌进了 Lucene,然后在不同的、合适的 section 下面做搜索。对。其实我们直接看一下正在跑的这个查询可能更好理解。这里就是……我把它放大一点。


[1:13:32]

Um so here's here's the cipher query and I'll I'll create a space here so you can kind of see it better. So we have a lucine index that we've set and we set it when we loaded the graph. We set it on the document and the section nodes. So every time that I call this index full text query nodes, it's going to hit the name of the index which is content search and then this parameter lucine is whatever I've fed it in with that

这就是那条 Cypher 查询,我加个空行让你们看得清楚一点。我们设置了一个 Lucene index,是在 load graph 的时候建的,建在 document 和 section 节点上。所以每次我调用这个 index 的 full text query nodes 时,它会命中这个 index 的名字,也就是 content search,然后这个 Lucene 参数就是我通过那个脚本


[1:14:03]

script. And so that's going to do our initial filter to just a bunch of documents and sections because we've structured our IDs for every node as a URI that's hierarchical. If I know that I only want to search underneath recall notices, I can hand it that URI and it will say only nodes whose URI starts with that thing. So it's applying that post filter afterwards if that makes sense.

喂进去的内容。这就完成了初步过滤,把范围缩到一批文档和 section。又因为我们把每个 node 的 ID 都设计成了层级化的 URI,如果我知道我只想在召回通知(recall notices)下面搜索,我就可以把那个 URI 传给它,它就会只返回 URI 以那个前缀开头的节点。所以它是在之后再套上这个后置过滤器,不知道这样讲清楚了没有。


[1:14:32]

Yeah. Yep. You have a question. So the question is if say you already have a graph and you haven't created your lucine index yet, could you prompt claude code to basically do this for you? Um so I I think yes to be honest I might have done that with this course initially. I might have actually had Claude do it. I can't remember because I' I've rebuilt this in in a in a few different ways. Um, but I think the answer to that is yes, your mileage may

对。你有问题。这个问题是,假设你已经有一个 graph,但还没建 Lucene index,你能不能让 Claude Code 帮你把这件事做了?说实话我觉得可以,我这门课最开始可能就是这么干的,我可能真的是让 Claude 做的。我记不太清了,因为我用好几种不同方式重建过这个东西。但我觉得答案是可以的,只是效果会


[1:15:21]

vary depending on your data. A lot of times with when you set up a lucine index too, it's like especially it's very flexible inside of Neo forj. So it's like you can set it up on top of one node or multiple nodes, multiple properties on one node, multiple properties between multiple nodes. So,

因你的数据而异。很多时候在建 Lucene index 的时候,特别是在 Neo4j 里,它非常灵活。你可以把它建在一个 node 上,也可以建在多个 node 上,可以是一个 node 的多个属性,也可以是多个 node 之间的多个属性。所以,


[1:15:39]

you know, there there's a lot that flexibility can also be a little bit of an Achilles heel too when you ask an AI model to do because it might not understand, you know, all those various options. But in general, yes, you you can have um you can use agentic coding to to help put this together for you.

你知道,这种灵活性在你让 AI 模型去做的时候,也可能成为它的阿喀琉斯之踵,因为它可能理解不了所有那些不同的选项。但总体来说,是的,你可以用 agentic coding 来帮你把这套东西搭起来。


[1:15:57]

Um yes, in the middle there So the question is how important is it that the name of the documents be accurate for the agent to traverse it effectively? Um,

中间那位请讲。这个问题是,文档的名字取得准不准确,对 agent 能否有效遍历有多重要?


[1:16:30]

I think I think it would. Yeah. If if your documents weren't named very well and you had that outline view, then the agent doesn't have as much to go off of when it's trying to traverse. So, it might misinterpret what a document means, right? Um that's why having the search augmentation alongside of it is useful because it can actually look in the document. Um with these things what I've seen because I use this actually

我觉得会有影响。对。如果你的文档命名很糟糕,而你又用那种大纲视图,那 agent 在尝试遍历的时候就没什么可依据的了。所以它可能会误解某个文档的含义,对吧?这就是为什么在旁边配一个搜索增强很有用,因为它可以真的去文档内部看一看。关于这些东西,根据我的经验——因为我自己实际上也用这套来做


[1:16:57]

for my own knowledgebased management. I have my own open source library that I use which has these tools. Um the biggest problem is when you have documents that are outdated or drifted, right? Um, so I've thought about, well, maybe we need to add like a last updated date or something or like um whether or not a document should be authoritative,

我个人的知识库管理。我有自己的开源库,里面就有这些工具。最大的问题是当你的文档已经过时或者发生了漂移的时候,对吧?所以我在想,也许我们需要加一个“最后更新日期”之类的字段,或者标记一个文档是不是权威版本,


[1:17:17]

right? Um, and the document naming is actually important. Um, and link naming as well is important too. Um, so if you're if you have links that um don't have like um I forget what it's called, but like in Obsidian, you can sort of you can name the you can give it um uh a synonym or something, right? where you name the link like that can be very helpful if that's there right because then it's like oh I know exactly why I'm

对吧?文档命名确实重要,链接命名同样也重要。所以如果你的链接没有……我忘了那个东西叫什么了,但比如在 Obsidian 里,你可以给链接起个别名或者同义词之类的,对吧?你给链接这样命名的话,是非常有帮助的,因为这样就是“哦,我立刻就知道我为什么要


[1:17:43]

linking out to this other thing um so that's when sometimes using like a language model inside of the ingest if you don't have that could could be beneficial um yes over here so when you get a new document that comes in um so the nice thing about this is that it's all an item potent load which means that say your entire graph went away tomorrow as long as your documents didn't change if you load them you'll get the same

链接到另一个东西上。所以这就是为什么有时候在 ingest 环节里用一下语言模型——如果你没有这些链接信息的话——可能会挺有帮助的。好,这边这位。所以当有新文档进来的时候……这里比较好的一点是,整个加载过程是幂等(idempotent)的,也就是说,假设你的整个 graph 明天全没了,只要你的文档没变,你重新加载一遍,得到的还是同一个


[1:18:20]

graph so worst case scenario yes you would have to reload the whole graph but this load is deterministic now you can also and I've experimented with this again in some of my other um work is you can say oh well I have one document that changed and I just want to update that one node. You just have to be aware of the fact that it links to other things.

graph。所以最坏的情况就是,是的,你得把整个 graph 重新加载一遍,但这个加载过程是确定性的。当然你也可以——我在其他一些工作里也试过——你可以说,哦,我只有一个文档变了,我只想更新那一个 node。你只需要意识到一点:它是会链接到别的东西的。


[1:18:42]

So it's like well if you got if you changed if you updated this document over here um if you changed its name then you have to you know did your other documents also update like the URL reference to that document. So there's little edge cases there that you have to work through. But um if you're able to do that maybe we can talk after and I can show you. You can have like just add to like this part of the tree and clean

所以就会变成这样:如果你更新了这边这个文档,如果你改了它的名字,那你得考虑,你其他文档里指向这个文档的 URL 引用是不是也要跟着更新。所以这里会有一些小的边界情况需要你去处理。但如果你能搞定这些,也许我们会后可以聊聊,我可以演示给你看。你可以只往树的这一部分里添加内容,然后清理


[1:19:08]

up this part of the tree so you can do like partial sinking. Yep. All righty. Um is do you have your hand up back there? You're just scratching your head. All right. Um then I guess we have one question here. right? the default after something like page index would be search.

树的那一部分,这样就能做到部分同步。对。好嘞。后面那位是举手了吗?哦你只是在挠头。好的。那我想我们这边还有一个问题,对吧?像 page index 这种做完之后,默认的下一步应该就是 search 了。


[1:20:11]

Yeah. Yeah. So I think the question is the question basically like um for the way that the agent reasons about it that it would call outline first to then use like search. Yeah. Yeah. It it could work that way. It could work some way where it will search first to find a relevant document and then try to find everything it links to. In which case,

嗯,对。我想这个问题基本上是在问,按照 agent 的推理方式,它会先调用 outline,然后再用 search。对,这样是可行的。也可能反过来:它先做 search 找到一个相关文档,然后再去找这个文档链接到的所有东西。那种情况下,


[1:20:32]

it would be the opposite way around. All right, awesome. Let's um let's move on here then to make sure that we have time for all of our questions. Um so what are we doing here? This is just more asking more full text search. Why don't we because we're at 39.

顺序就是反过来的。好的,很棒。那我们继续往下走吧,确保有时间回答大家所有的问题。那我们这里在做什么呢?这部分基本上就是更多的全文搜索(full text search)内容。要不我们……因为现在已经 39 分了。


[1:21:02]

Well, I think we already effectively went over a lot of this material here. Um, some of this is designed so that if you come back to it later, it kind of overdocuments what I'm talking about. So these steps and some of the optional work will help you understand cipher a little bit better. Um and kind of exactly how exactly how all the pieces fit together. But given given the pace that we're moving at, why don't I go

嗯,我觉得这部分材料我们实际上已经讲了很多了。有些内容的设计初衷是,如果你之后再回来看,它会把我讲的东西过度文档化一遍。所以这些步骤和一些可选的练习能帮你更好地理解 Cypher,以及所有这些部件到底是怎么拼在一起的。但考虑到我们现在的推进速度,要不我


[1:21:30]

ahead and jump over to our um theme section. So I'm going to skip over the optional practice lesson here and I'm going to go right into into themes. So this is going to be our third and uh last shape of today. The the idea with this is if you run so this cipher query is basically going to be loading a graph and I'm just looking at that links to relationship between the documents. So when you have documents that have a lot of

直接跳到我们的 theme(主题)这一节吧。所以我要跳过这里的可选练习课,直接进入 themes。这将是我们今天讲的第三个、也是最后一个 shape。它的思路是这样的:这个 Cypher 查询基本上是在加载一个 graph,我这里只看文档之间的 links to 这个 relationship。所以当你的文档之间有大量的


[1:22:05]

interlinking um it's always nice to think about using a graph because this structure basically what it's saying here is right we have these manuals these recalls and these bulletins and they link and they refer to each other. Um, so you can see like this, you know, um,

互相链接时,用 graph 来思考总是很不错的,因为这个结构本质上表达的是:我们有这些手册、这些召回通告、这些技术公告,它们彼此链接、互相引用。所以你能看到,比如说,嗯……


[1:22:26]

uh, this this this guide here that's a um that's a manual is being linked to by all of these different repair procedures. And if if you zoom out, what ends up happening is you'll get natural clusters of things. And this also happens a lot like in these kaparthy style like knowledge bases where you'll see that like concepts will naturally start grouping together and a graph can help you surface those themes those things that

呃,这里这份指南,它是一份手册,被所有这些不同的维修流程文档链接着。如果你把视野拉远,最终会发现事物自然形成了聚类。这种情况在那些 Karpathy 风格的知识库里也经常发生——你会看到概念自然而然地开始聚到一起,而 graph 能帮你把这些主题浮现出来,也就是那些


[1:22:57]

you didn't actually know existed before. And so we're going to use something called lien community detection. By a show of hands, how many people in this room are familiar with what graph data science is? Well, yes. Okay. So, so not too many of you. Do how many people know what community detection in a graph is?

你原本并不知道存在的东西。所以我们要用一个叫做 Leiden 社区检测(community detection)的算法。举手示意一下,在座有多少人了解什么是图数据科学(graph data science)?嗯,好的。看来没有太多人。那有多少人知道 graph 里的社区检测是什么?


[1:23:21]

Okay. So, we we have some of you. So, so community detection is this idea where right if if I have this graph um and you can kind of see if I zoom out if you if especially if you're running it locally you can see that there's these like natural little like clusters right of nodes that are highly interlin together and so the idea is like well what if we can like try to label these clusters such that within a cluster things are highly interconnected so like

好,有一些人知道。那么,社区检测的思路是这样的:如果我有这么一个 graph,你把视野拉远——尤其是如果你在本地跑的话——你能看到有这些天然的小簇,对吧,就是一堆彼此高度互链的 node。所以思路就是:我们能不能想办法给这些簇打上标签,使得同一个簇内部的东西高度互联。比如说


[1:23:50]

this little globule of nodes becomes a cluster cluster, then this globule becomes a cluster and this one down here. Um, and and if if we can do that, we can start understanding our data at a global scale really well. So, this is like that initial like Microsoft graph rag idea too of like of global versus local search, right? We're doing something very similar here, but we're doing it in a very lightweight way where we're not really doing any LLM

这一小坨 node 变成一个簇,那一坨变成另一个簇,下面这个再一个。如果我们能做到这一点,就能在全局尺度上很好地理解我们的数据了。这其实也很像最初微软那套 GraphRAG 的思路——全局搜索 vs 局部搜索,对吧?我们这里在做非常类似的事情,但方式非常轻量,我们基本上没有用任何 LLM


[1:24:15]

extraction. We're just going off of literally the the structure of the documents. The algorithm that we're going to use for that labeling is called Leiden, which is similar to Louane, which if you talk about community detection, Luane will come up a lot.

抽取。我们完全是基于文档本身的结构。我们用来做这个标注的算法叫 Leiden,它和 Louvain 很像——如果你聊社区检测,Louvain 会经常被提到。


[1:24:30]

Leiden is I almost see it as like a um the sort of next step. It's a little bit more efficient um in in the way that it runs. And um it it's using Neo Forj's graph data science library. So basically what we have to do or what what we do to make it very performant is in addition to just having the database we have this other um projection where we'll take a part of the graph into memory um and then we'll run these high concurrency

我基本上把 Leiden 看作是下一代的版本。它在运行方式上效率更高一些。它用的是 Neo4j 的图数据科学(graph data science)库。所以我们要做的,或者说为了让它跑得非常高效我们所做的,就是在数据库之外,我们还有一个 projection(投影),我们把 graph 的一部分放进内存,然后在上面跑这些高并发的


[1:25:03]

algorithms on top of that so that if you have a graph that has say millions or billions of nodes we can start performing this clustering and then identifying um basically the different interconnected communities. and um these clusters. If I um keep going down, basically the the format that we're going to go for here,

算法。这样,即使你的 graph 有几百万甚至几十亿个 node,我们也能开始做这种聚类,然后识别出不同的互联社区,也就是这些簇。我再往下翻,我们这里要采用的格式基本上是这样,


[1:25:26]

sort of the view that we're going to show our agent is this one. And so it's a little bit hard to see because it's a sliding window, but basically we'll have what are what we're calling themes. And these themes are going to be these buckets of documents. we'll have a sense of how tightly or loosely they're interlin using um this conduance metric.

我们要展示给 agent 看的视图就是这个。它有点不太好看清,因为是个滑动窗口,但基本上我们会有所谓的 themes(主题)。这些 theme 就是一个个文档的桶。我们会知道它们之间互链得有多紧密或多松散,用的是 conductance(电导)这个指标。


[1:25:46]

It's basically like going to be this metric around like how interconnected the nodes inside cluster are versus how much they um are connected outside of the cluster. So we can say something's tightly interlin or loosely interlin. um we'll use the labels on the links as the top shared targets and we will um talk about the most linked docs and so so the highest centrality docs inside each of them and so what we end up with

这个指标基本上衡量的是:簇内部的 node 之间连接有多紧密,相对于它们向簇外部连接的程度。所以我们可以说某个东西是紧密互链还是松散互链的。我们会用链接上的标签作为 top shared targets(最常被共享的目标),我们还会展示被链接最多的文档,也就是每个簇里中心度最高的文档。所以我们最终得到的结果,


[1:26:16]

without any sort of um tagging or labeling by AI or a language model um is just simply from the document structure we can tell that this first one um is about you know um BCM um and and bus and all these sorts of things. And then if if I go down like this next cluster is going to be about brakes um and rotor pads and um hydraulic lines. So it's all braking stuff. So you end up with these very natural sort of clusters that come up.

在完全没有用 AI 或语言模型做任何打标、标注的情况下,仅仅依靠文档结构,我们就能看出第一个簇是关于 BCM、总线(bus)之类的东西。然后我再往下翻,下一个簇是关于刹车的——刹车盘、刹车片、液压管路。所以全都是制动相关的东西。所以你最终得到的是这些非常自然浮现出来的聚类。


[1:26:51]

And so this is very useful from sort of a whole estate wide question because you can start to understand in your data kind of how everything kind of groups and clusters together. Uh and so the way that we build that is similar to what we were doing before where we have that script and we have our um our spec as well. Um,

这对于回答那种全域(estate-wide)范围的问题非常有用,因为你可以开始理解你的数据整体上是怎么分组、怎么聚类的。我们构建它的方式跟之前做的类似:我们有那个脚本,也有我们的 spec(规格说明)。嗯,


[1:27:17]

so collapsing the sections and it talks a little bit here and I I kind of want to give us the last half an hour to go through the final questions. So I'll speedrun this a little bit, but basically when we create this projection and I'll just go ahead and and copy this thing. Maybe I'll I'll do this script thing first. So like we did before, um,

我把这些章节折叠起来,这里有一些说明文字。我想把最后半小时留出来过一遍最终的提问环节,所以这部分我会稍微速通一下。基本上,当我们创建这个 projection 的时候——我先把这个东西复制一下。或者我先来搞这个脚本吧。跟我们之前做的一样,


[1:27:39]

inside of here, uh, if I look inside of my themes.py Pi file I I have this wire for this for the projection that piece where we take the graph and put it into memory. Um the actual liiden algorithm here I'm just calling it with that graph data science library. Um so in theory you can you can kind of agent code all of these. Um but I'll just do this one. Um oh that's not good. I want to be able to make sure I copy the right thing. So, similar to

在这里面,如果我看一下我的 themes.py 文件,我这里有为 projection 准备好的接线代码,就是把 graph 放进内存的那一部分。至于真正的 Leiden 算法,我这里就是直接调用那个图数据科学库。理论上,这些你都可以让 agent 帮你写代码。不过我就演示这一个吧。哦,这个不太行。我得确保我复制的是对的东西。所以,跟


[1:28:08]

before, um, we have our, um, sure that the right thing that I just gave it. Yeah. So, we're telling it to use a cipher skill and the GDS skill to put this together. If you went inside of our docs, you have the uh, theme format one, which will give it the the shape that we want and everything. Um, and it will tell us kind of how we want all the headers and things to be formatted.

之前类似,我们有我们的……好,确认一下我刚才给它的是对的东西。对。所以我们告诉它使用 Cypher skill 和 GDS skill 来完成这件事。如果你进到我们的 docs 里,你会看到 theme format 这份文档,它会告诉模型我们想要的 shape 之类的一切信息,也会告诉我们所有的表头之类的东西该怎么格式化。


[1:28:35]

Um, and so we'll go ahead and um create that create that for me. And while it's creating that, I'll talk a little bit about the projection. So there's some graph manipulation that we do to bring it efficiently into a projection. Um, basically we collapse the relationships on the URI so that the if we have sections that interlink with each other like one section of a doc interlinks to another subsection, we just aggregate

好,那我们就让它帮我创建出来。趁它在生成的时候,我来讲讲 projection 这块。我们会做一些 graph 的变换处理,好让它高效地进入 projection。基本上我们会按 URI 把 relationship 折叠合并,这样如果我们有一些相互链接的 section,比如一个文档的某个 section 链接到另一个子 section,我们就把这些全都聚合


[1:29:07]

that all to the document level. Um, and that creates kind of a cleaner interpretation to how our documents link together rather than just all our individual sections and we can get out the communities um, easier that way as well. All righty. So, that looks like it's coming along.

到文档层级上。这样对于我们的文档之间如何链接,就有了一个更清晰的解读,而不是全部停留在一个个单独的 section 上,而且这样我们也更容易提取出社区。好嘞。看起来这边已经差不多了。


[1:29:34]

Yes. Okay. So, it went ahead and created that for us. Um, and then once we get that back, uh, I can go ahead and copy the script that it just helped me create and I can run that inside of my terminal. Let me just make a new terminal. So, I can go ahead and and call that theme script that we just edited. And then it will do that community detection. And you'll see here that I'll get that view. Um so I have my you can

好,它已经帮我们创建好了。然后拿到结果之后,我可以把它刚帮我生成的脚本复制过来,然后在终端里运行。我新开一个终端。好,我可以运行我们刚编辑好的那个 theme 脚本。然后它就会去做社区检测。你们会看到我这里能得到那个视图。嗯,你们可以


[1:30:11]

see like the this you know circuit you know one I have all my different uh documents that are that are grouped together and I kind of understand now um all the different areas of the vehicles that that my documents go over. Um there's other parameters that you can feed this. Um so for example we have a gamma parameter and basically what that will do is uh basically tells it kind of um how how much to split everything

看到,比如这个电路相关的这一组,我所有不同的文档都被归到了一起。现在我大概就能理解我这些文档覆盖了车辆的哪些不同领域。这里还有其他参数可以传给它。比如我们有一个 gamma 参数,它的作用基本上是控制把所有东西切分


[1:30:41]

into. So it's one parameter that we're exposing. If you went to the graph data science documentation there's many others. Um but like I got 13 groups running it here. If I if I try to make it more refined with gamma equal to two. So it's one by default. it will give me I think in this case 14. Um so it basically splits it up into more groups. A lot of these community detection algorithms also because they're hierarchical you can choose to

得有多细。这是我们暴露出来的一个参数。如果你去看图数据科学的文档,还有很多其他参数。这次跑下来我得到了 13 个组。如果我把 gamma 设成 2、让它更细一点——默认是 1——这种情况下我记得会得到 14 个。所以它基本上就是把结果切得更碎、分成更多组。另外,很多这类社区检测算法因为是层次化的,你还可以


[1:31:11]

sort of you can choose your level of granularity that you want. Basically um as you go along and you start to understand your corpus better, you may want to tune some of these hyperparameters. Um and this gave us 14 different themes. Um, let's see what we have here.

选择你想要的粒度层级。基本上,随着你不断推进、对自己的语料库理解得越来越深,你可能会想去调这些超参数。这次给了我们 14 个不同的 theme。我们来看看都有些什么。


[1:31:36]

Yeah. And it never really names a theme is sort of the point of this. So like everything that you get um is just directly from the document. If I was to look at like the the theme here, the names of the links, the the names of the the top files that were read, all of those things are just directly from the data. And then because you get the URIs,

对。注意,它其实从来没有给 theme 起过名字,这正是这套做法的关键点。所以你拿到的所有东西都是直接来自文档本身。如果我看这里的这个 theme,链接的名称、读取到的那些排名靠前的文件的名称,所有这些都是直接从数据里来的。然后因为你能拿到 URI、


[1:32:01]

the IDs of some of these documents, you can start doing the thing where you can go back to the outline shape or the search shape and you can like search underneath all of these things. So, it gives the agent the ability to kind of jump between multiple tools like this.

拿到这些文档的 ID,你就可以做接下来的事情:回到 outline shape 或者 search shape,在这些东西下面继续做搜索。所以这就给了 agent 在多个工具之间来回跳转的能力。


[1:32:18]

Um and so this the these sorts of things are good for these estate level questions. So things like where are issues concentrated? Where's documentation thing thin? Where does new documentation possibly belong on some other different type of subject that just came in for adding more data.

所以这类东西非常适合回答全域(estate 级别)的问题。比如:问题都集中在哪些地方?文档在哪里比较薄弱?如果新来了一批关于某个不同主题的数据,新的文档应该归到哪里去?


[1:32:36]

Um all righty. So that brings us to the end of themes and we have 27 minutes about left. Are there any questions around this themes uh algorithm that anyone has really quickly? Uh yes.

好嘞。这样 themes 这一部分我们就讲完了,我们大概还剩 27 分钟。关于 themes 这个算法,大家还有什么问题吗?快速问一下。好,请说。


[1:32:57]

Yep. Yeah. So that's a good question and the question is once you get the themes so all the line algorithm will do is assign an ID to these different groups right and then what you do after that is sort of your choice and your question is are you just sort of taking from the data and just putting that there or are you making some sort of inference afterward to label the different groups. Here we're doing the former. We're just

嗯。对,这是个好问题。问题是:当你拿到这些 theme 之后——Leiden 算法所做的全部工作就是给这些不同的组分配一个 ID,对吧——那之后你要怎么做,其实取决于你自己。你的问题是:我们是只是从数据里取出内容直接放在那儿,还是之后再做某种推断来给不同的组打上标签?我们这里用的是前者。我们只是


[1:33:53]

taking what's in the data and we're just showing it to you. Um, and the advantage of that is every time I run it, it'll be the same as long as the data stays the same. If the data changes, it will change to reflect the data. So, it's very stable. The disadvantage to it would be if your links in your documents um, and the titles and things that are being scooped up, because this is really only looking basically at document

把数据里已有的内容取出来展示给你。这样做的好处是,只要数据不变,我每次跑出来的结果都是一样的。如果数据变了,结果也会随之变化以反映数据。所以它非常稳定。它的劣势在于,如果你文档里的这些链接、标题以及被抓取出来的这些东西——因为这套方法基本上只看文档


[1:34:18]

metadata and link metadata. um if those things aren't already wellleeled, this view might not be super informative right off the bat. That's why you have, you know, with like the the graph rag methods that Microsoft came up with, they do a lot of heavy entity extraction because then what that will give is this sort of and they'll do um hierarchical level summaries, right? So after they get the light in community, which is the

元数据和链接元数据。嗯,如果这些东西本身没有被很好地打好标签,那这个视图一开始可能不会特别有信息量。这就是为什么,比如说微软提出的那套 GraphRAG 方法里,他们会做大量的 entity extraction,因为这样能得到那种……而且他们还会做分层级的摘要,对吧?所以在他们跑完 Leiden 社区发现之后——就是我们用的


[1:34:44]

same algorithm that they use, they'll do a summary that'll be LLM driven on top of each theme. Um, so that's possible to do. Um, but then obviously it costs more money, it's slower, and the you if you run it twice, it might not return the same thing. So there's trade-offs between uh each way of doing it. I'm showing you the sort of lighter way of doing it and the easier way. Um, just because it's faster and if you're just

同一个算法——他们会在每个主题之上再做一层由 LLM 驱动的摘要。嗯,这是可以做的。但很显然,那样成本更高、速度更慢,而且你跑两遍,结果可能不一样。所以每种做法之间都有取舍。我给大家展示的是比较轻量、比较简单的做法。嗯,就是因为它更快,而且如果你才刚


[1:35:10]

getting started with this, it might be easier to to start there. Um, may we have time for maybe one maybe two more questions. Um, yes.

开始接触这块,从这里入手可能更容易上手。嗯,我们可能还有时间再回答一个、也许两个问题。嗯,你请说。


[1:35:25]

Temporal data.

时序数据(temporal data)。


[1:35:31]

Yeah. I mean, um, I've seen customers, you know, basically every time because because the way Leiden works is it's it really is this algorithm that where it will suck everything into a projection. it will create its labels and it'll it'll die down and then it will go away. So if your data is being updated constantly um you can recreate your ids um and you can have almost like a time series of of different theme ids for example um and

对。我是说,嗯,我见过一些客户,基本上每次都……因为 Leiden 的工作方式是,它确实是这么一个算法:它会把所有东西吸进一个投影里,生成它的标签,然后收敛下来,接着这个投影就销毁了。所以如果你的数据在不断更新,嗯,你可以重新生成这些 id,嗯,你就可以得到差不多像一个时间序列一样的、不同时刻的主题 id,举个例子,嗯,然后


[1:36:03]

then you can if you want create summaries sort of at the snapshot of when they existed or you can see how they evolved over time um but it's a very common use case a lot of people will actually use this even before AI they'll use this for things like um fraud detection so like looking at um like credit chargeback fraud or like some of these other like um anti-moneylaundering type of stuff to look at clusters and for that they'll

如果你愿意,你可以针对它们存在的那个快照时点生成摘要,或者你可以观察它们随时间是怎么演化的。嗯,但这是一个非常常见的用例,很多人其实在 AI 出现之前就已经在用了,他们会把它用在比如说欺诈检测上,像是看信用卡拒付欺诈,或者其他一些像反洗钱那类的场景,用来看聚簇。而在这些场景里,他们


[1:36:28]

have to do it temporally like they have to keep running it um to kind of see how things change over time and then predict the future. Right. So they're used in those scenarios too. Yep. All right. One more maybe and then I'm gonna have to move on.

就必须按时间维度来做,他们得反复跑,嗯,来看事情随时间怎么变化,然后预测未来。对吧。所以在那些场景里也会用到。嗯。好的。再回答一个问题吧,然后我就得往下讲了。


[1:36:45]

All right. I Okay. Yes. So the question is if you have a really big graph, how do you take this view that's being sent to the AI model and make it manageable? Yeah. Um so for this view I think it's showing like 14 themes. So there are cuto offs that you can make. So for example you can say like for communities that are smaller than X threshold like you don't necessarily need to highlight them. Sometimes small communities are

好的。我……好,你请说。所以问题是:如果你的 graph 非常大,你怎么把这个要发给 AI 模型的视图控制在可管理的规模?对。嗯,就这个视图来说,我记得它显示了大概 14 个主题。所以你是可以设一些截断阈值的。比如说你可以规定,对于小于某个阈值 X 的社区,你就不一定需要把它们凸显出来。当然有时候小社区反而


[1:37:34]

very important which is why that conduance metric which is powering that tightly interlin loosely inter interlin you can also make that a cutoff. So if I'm interpreting your question correctly, it's sort of filtering down kind of the amount of information to what's most important to show the AI model on a larger graph. Do I understand that correctly?

非常重要,这也正是那个 conductance 指标的作用——就是驱动「紧密互联/松散互联」判断的那个指标——你同样可以拿它来做截断。所以如果我没理解错你的问题的话,你问的其实是:在更大的 graph 上,怎么把信息量过滤收敛到最重要的那部分再拿给 AI 模型看。我理解得对吗?


[1:38:00]

to make sure that the relationships on the nodes are correct. And um so yeah. Yeah. I mean, so in this scenario, we're really sort of trusting, we're taking the source data kind of at its word, right? If something links to a document, it links to the document,

……要确保 node 上的 relationship 是正确的。嗯,所以,对。对。我是说,在这个场景里,我们其实是选择相信,我们基本上是照单全收地采信源数据的,对吧?如果某个东西链接到一份文档,那它就是链接到那份文档,


[1:38:36]

right? If it's erroneously linking to a document or there's a section that's malformatted, um, we wouldn't necessarily catch that. But I suppose a good thing about something like an outline shape is that you can have your agent traverse it um automatically without you necessarily seeing it in small pieces and it might be able to catch some of those things,

对吧?如果它是错误地链接到某份文档,或者某个章节格式是坏的,嗯,我们不一定能发现。不过我想,像 outline 这种形状的一个好处是,你可以让你的 agent 自动去遍历它,而不需要你自己一小块一小块地去看,这样它也许就能发现其中一些问题,


[1:38:57]

right? Um so I suppose it does give you that navigation would give you a way to kind of have an agent supervise the graph and understand like malformed data or data that's been a problem. Um but it's a good question. I don't know if we have a perfect solution to it. It's cleaning and cleaning messy data has always been a, you know, a thing. Yep.

对吧?嗯,所以我想它确实给了你这种导航能力,让你能用一个 agent 去监督这个 graph,识别出格式错误的数据、或者一直有问题的数据。嗯,但这是个好问题。我不知道我们有没有一个完美的解法。清洗、清洗脏数据一直都是个老大难,你懂的。嗯。


[1:39:20]

All righty. Let's go ahead and and move on because I only have 20 minutes left and I anticipate this last uh section will be this is the the meatiest one. So hopefully we'll have enough time to go through it especially because Claude has been slow. So cross your fingers um because this one is probably the heavier use of claude. So there is a section here around just using the Neo Forj CLY because we're just 20 minutes in. Um I'm

好嘞。我们继续往下走吧,因为我只剩 20 分钟了,而且我预计最后这一部分会是……这是最有干货的一部分。所以希望我们有足够的时间讲完,尤其是因为 Claude 一直有点慢。所以大家祈祷一下,嗯,因为这一部分对 Claude 的使用量可能是最重的。这里有一节是讲怎么用 Neo4j CLI 的,因为我们才进行到 20 分钟。嗯,我


[1:39:46]

probably end I'm not going to go through it all the way but I'll talk about it for a couple minutes. The Neo Forj CLY is a a Cly tool. So I can run it in the command line. So here for example um if I opened up a terminal window I can go ahead and copy it in and it will run a uh query for me and it also has the ability here to grab the graph schema.

可能不会……我不打算把它完整讲一遍,但我会花几分钟聊一下。Neo4j CLI 是一个命令行工具。所以我可以在命令行里跑它。比如说这里,嗯,如果我打开一个终端窗口,我可以把它复制进去,它就会帮我跑一个查询,而且它这里还有能力去抓取 graph schema。


[1:40:11]

Um and what that enables me to do is um basically understand what's in the graph and then write a query based on that. And if you have your agent which in this case because we're using a coden agent it can access the neo forj cli it gives it the ability to do this graph reasoning read the schema and then do flexible queries. So this is very useful if you have a question that you didn't anticipate and the agent's sort of

嗯,这让我能做的事情就是,嗯,基本上先搞清楚 graph 里有什么,然后基于这个去写查询。而如果你有一个 agent——在这个例子里因为我们用的是编码类 agent——它可以访问 Neo4j CLI,这就赋予了它做 graph 推理的能力:读取 schema,然后做灵活的查询。所以当你有一个事先没预料到的问题时,这就非常有用,agent 就得自己


[1:40:37]

gluing things together between the shapes, right? It has to write its own custom cipher query. It can do that very efficiently. Um, and I've seen a lot of improvements using this along with the skills. So much better than the text to cipher experience that we've had like even as as soon as a year ago or six months ago. Um, if you are going to be writing cipher or doing anything text to cipher with an agent, I'd highly

在这些 shape 之间把东西拼接起来,对吧?它必须自己写一个自定义的 Cypher 查询。它可以非常高效地做到这一点。嗯,我看到把它和 skills 一起用之后有很多提升。比起我们哪怕一年前、或者半年前的 text-to-Cypher 体验要好太多了。嗯,如果你打算写 Cypher,或者要用 agent 做任何 text-to-Cypher 的事情,我强烈


[1:41:01]

recommend using the Neo Forj CLY. Um, as well as the cipher and GDS skills that we were just going over. So, with that in mind, I'll go ahead and take it on to our, in this case, our last section. If you were to take this offline, there's otherformational sections that come after it.

推荐使用 Neo4j CLI。嗯,以及我们刚刚过了一遍的那些 Cypher 和 GDS skills。那么带着这些,我就直接进入我们的——在这个场景里是——最后一节了。如果你把这些材料带回去自己看,后面还有一些其他的信息性章节。


[1:41:21]

Um but basically we're going to start asking um some questions. Uh the first question that we're going to ask it and I will jump right to it because a lot of this documentation we've already went over is if you remember from the beginning we had our different personas right we had sort of our our who we called our Danny which is the floor technician right and they might have a question like hey for this VIN with this

嗯,但基本上我们要开始问一些问题了。呃,我们要问的第一个问题——我会直接跳到那里,因为里面很多文档我们已经过过了——如果你还记得开头的话,我们有几个不同的 persona,对吧?我们有那个我们叫 Danny 的角色,就是车间技师,对吧?他可能会有这样的问题:嘿,对于这个 VIN、报出这个


[1:41:50]

specific code that I'm getting what fix what have we done that has fixed this on similar vehicles. Um, and the motivation behind this, right, and behind a lot of the data is as an auto repair shop, you want to minimize your um, comeback ratio, which is basically how many times a customer has to come back because, you know, the fix didn't work, right? Um,

特定故障码的情况,我们在类似车型上做过哪些修复动作是有效的?嗯,这背后的动机,对吧,也是很多数据背后的动机是:作为一家汽车维修店,你希望把返修率降到最低——返修率基本上就是客户因为修完没解决问题而不得不再跑一趟的次数,对吧?嗯,


[1:42:16]

and so the data, especially on the warehouse side, will show um some of that history combined with the documentation on the actual parts and the recalls and the bulletins. So, what I've done here is similar to the gentleman's questions before, I uh I have it explained the steps that it used for the different shapes. You can see it's loading um the auto server skill.

所以这些数据,尤其是数仓那一侧的数据,会展示出一部分这样的历史记录,再结合上关于实际零部件、召回和技术通报的文档。所以我在这里做的,跟前面那位先生问的问题类似,我让它把针对不同 shape 所用的步骤解释出来。你可以看到它正在加载,嗯,那个 auto server skill。


[1:42:40]

So that was the um the skill that we have here where we've also contains all of the scripts we've been working with. So the outline search and theme script it has access to as well as the run SQL and the MCP server. Um that was this skill that we were talking about before.

那就是我们这里的那个 skill,我们把一直在用的所有脚本也都放在里面了。所以 outline search 和 theme 脚本它都能访问,还有 run SQL 和那个 MCP server。嗯,就是我们之前聊过的那个 skill。


[1:42:58]

So I've pre-written that for this. You can read it um if you like, but it basically gives it um some general guidance on how to access the warehouse and also deal with the different shapes in the command interface. And you'll see what it will do here. Um it will look for the document code.

所以我为此提前把它写好了。你可以读一读,嗯,如果你想的话,但它基本上就是给它一些总体指引,告诉它怎么访问数仓,以及在命令行界面里怎么处理这些不同的 shape。你会看到它接下来会做什么。嗯,它会去找那个文档编号。


[1:43:19]

It'll do a full text search. It'll use the tree shape here to find cross links and causes. Um and then it will get the join paths. Um, and it will actually do the query for the VIN. Um, and when it does that, it will go ahead and come back with basically the part number that needs to be replaced. And in this case,

它会做一次全文检索。它会用这里的 tree shape 去找交叉链接和成因。嗯,然后它会拿到 join path。嗯,接着它真的会针对这个 VIN 做查询。嗯,做完之后,它就会返回基本上是需要更换的那个零件号。在这个案例里,


[1:43:43]

there was like an old ignition coil that got revised that it had to replace it with. Um, so if you look right, it at first it did full text search to sort of ground uh the document with the right code. Um so it looked for the code and also the misfire or rough idle. Um it found uh the uh top hit which is this engine type. It confirmed its grounding.

有个旧款点火线圈后来做过改版,必须用新版来替换。嗯,所以你看,对吧,它一开始先做了全文检索,把文档和正确的故障码对上、做 grounding。嗯,所以它去搜了那个故障码,还有 misfire(失火)或者 rough idle(怠速抖动)。嗯,它找到了,呃,命中率最高的那条,也就是这个发动机型号。它确认了自己的 grounding。


[1:44:09]

So it shows the links from that uh manual over to these different procedures. And then from there it did a join path. Um it basically looked at all the work orders um with that DTC code uh for that part. Um and then it was able to bring back the ultimate question which is like you know from all the parts that were replaced and how we dealt with that code it went to the warehouse to grab that information. So you can see like this is a simple

所以它展示了从那份手册指向这些不同维修流程的链接。然后从那里它做了一次 join path。嗯,它基本上是把带有那个 DTC 故障码的所有工单都看了一遍,呃,针对那个零件。嗯,然后它就能够把最终那个问题的答案带回来——就是说,从所有被更换过的零件、以及我们过去是怎么处理这个故障码的角度,它去数仓里把这些信息抓了回来。所以你可以看到,这算是一个简单的


[1:44:37]

question. So if you were using vector search and like Genie, like their AI search and Genie and like data bricks, you could do this, but what often happens is to to basically find what it would need to do that linking on that tree shape, it would have to do much more vector hits,

问题。所以如果你用 vector search,再加上像 Genie 那样的东西——就是 Databricks 里的 AI search 和 Genie——你也能做到这件事,但经常发生的情况是,为了找到它在那个 tree shape 上做关联所需要的东西,它得做多得多的 vector 命中,


[1:44:55]

which each of those is is a chance, right, for this to run into an issue and not find the right document or potentially hallucinate and get something on a misfire that maybe wasn't n't related to that specific part, but because it grounded in the tree shape and everything, it was able to to get the right information and then link that back to the table.

而每一次命中都是一次机会,对吧,一次出问题的机会:找不到正确的文档,或者干脆产生幻觉,扯出某个跟这个特定零件其实无关的失火问题。但因为它是基于 tree shape 之类的东西做了 grounding,它就能拿到正确的信息,然后再把它关联回表里。


[1:45:17]

Um, and so it can be helpful and it can help with efficiency in these smaller questions. Um, but then what can happen is when you get to the estate level questions is where it can get really interesting. So, I'm going to go ahead and just copy this question and I'm going to let it start running. Um it seems like claude is moving faster now which is good but I'll talk about it as it's running. So this is a question around hey are there mismatches between

嗯,所以这是有帮助的,在这类比较小的问题上它能提升效率。嗯,但接下来会发生的是,当你面对全域级别(estate level)的问题时,才是真正有意思的地方。所以我要直接把这个问题复制过来,让它开始跑。嗯,看起来 Claude 现在跑得快一些了,这挺好,不过我会在它跑的时候继续讲。这个问题是关于:嘿,我们的


[1:45:45]

our documented procedures and problems that we're seeing in the field and what documents are missing or sort of on the other half like what documentations are we not leveraging at all inside of our warehouse data right because we have our documents which tell us about like the recalls and the bulletins and and all this sort of stuff um and the manuals but then we also have the work order history from our warehouse and so this

书面维修流程和我们在一线实际看到的问题之间,是否存在不匹配?哪些文档是缺失的?或者反过来那半边:在我们的数仓数据里,有哪些文档我们压根没有用上?对吧,因为我们一边有文档,告诉我们召回、技术通报之类的所有这些东西,嗯,还有手册;但另一边我们还有来自数仓的工单历史。所以这


[1:46:10]

is a question that someone in a supervisor role or someone in an analytics role might be interested in, right? Because it's sort and it's sort of like proving a negative or a mismatch because you're you're sort of saying like, you know, I I don't know what I'm looking for. I'm looking for a gap though. And you can imagine that with a tool like vector search, this would be a very hard question because sim or any type of similarity or lexical search

是一个做主管的人、或者做分析岗的人会感兴趣的问题,对吧?因为它有点像是……有点像在证明一个否定命题,或者说找一处不匹配,因为你其实是在说:你懂的,我也不知道我要找的是什么,但我要找的是一个缺口。而你可以想象,用 vector search 这类工具的话,这会是一个非常难的问题,因为相似度……或者任何类型的相似度检索或词法检索,


[1:46:37]

because like if you're just doing that alone by definition, you're search you can't really search for a negative. You have to search for things that are there. Um, and so what I found and I think what we found as a company is that graphs can be very useful when you start having these more global types of questions, these estate level questions that you want to ask, particularly if they're around patterns where you might

因为如果你只靠那个的话,按定义你根本没法去搜索一个「不存在」。你只能搜索那些确实存在的东西。嗯,所以我的发现,我想也是我们作为一家公司的发现是:当你开始问这些更全局性的问题、这些全域级别的问题时,graph 会非常有用,尤其是当这些问题是关于某种模式,而你甚至可能


[1:46:59]

not even know what you're looking for yet. Yes. And so this one takes a while to run because it at at a certain point here, it does have to take a large number of codes and join some data together. Um, but you'll see at the end here, it'll it'll kind of uh trickle in and it it'll it'll tell you how it went about finding everything.

还不知道自己在找什么的时候。是的。所以这个问题要跑一会儿,因为它跑到某个环节时确实得处理大量的故障码,还要把一些数据 join 起来。嗯,不过你会在最后看到,它会,呃,一点一点地把结果吐出来,然后告诉你它是怎么一步步找到所有这些东西的。


[1:47:24]

Um, and oftent times what you see with these, it'll it'll use some of the document data with some shapes and then it will go back and it will go back to the warehouse and query from there. um run through a lot of different things here. Um, another thing just to just while it's working on that, I did tell it here for sake of clarity to not use the near forjly. And the reason that I I have that here is so that you can see it

嗯,而且这类问题你经常会看到的是,它会先用一部分文档数据配合一些 shape,然后再回到数仓那边去查。嗯,在这里跑一大堆不同的东西。嗯,还有一件事,就趁它在跑的时候说一下:为了讲解清楚,我在这里特意告诉它不要用 Neo4j CLI。我这么设置的原因是想让你们能看到它


[1:47:58]

using the um the different shapes. Um, just for sake of understanding how the different shapes kind of fit together and work. Once you hand your agent the Neo Forj CLY um it becomes very powerful because it can start writing custom cipher queries and there will be instances where it'll prefer doing that over some prefix shape. Um, and I'd say that as our sort of Texas cipher capabilities and as we keep building up more skills, it'll start to prefer more

在使用这些不同的 shape。嗯,纯粹是为了让大家理解这些不同的 shape 是怎么组合、怎么协同工作的。一旦你把 Neo4j CLI 交给你的 agent,嗯,它就会变得非常强大,因为它可以开始写自定义的 Cypher 查询,而且会出现一些情况:它宁愿这么干,也不用某个预设的 shape。嗯,我想说的是,随着我们的 text-to-Cypher 能力提升、随着我们不断积累更多 skills,它会越来越倾向于


[1:48:28]

free form Neo forj Cly stuff uh much more frequently. And this one can sometimes take a little while to work through. So, I'll give it a another 20 seconds or so. Or maybe I can even open it up to a few questions while we're here. While we're waiting for this one. Yep.

更自由形式的 Neo4j CLI 用法,频率会高得多。而这个问题有时候要花点时间才能跑完。所以我再给它 20 秒左右。或者趁我们在这儿等的时候,我甚至可以再开放几个问题。就趁我们等这个跑完。嗯,你请说。


[1:49:02]

So, basically what it's going to do is it's going to look at the um well, I actually have to have it come back and remind me exactly how it um how it goes through. Um but basically what you can see from the work orders what has been worked on right and then you can sort of take the codes from there and see what documents it has been using. Um and then there's also going to be high comeback history on using some of the wrong parts. And so you sort of

所以它基本上会去看……嗯,其实我得让它回过头来提醒我它具体是怎么走完整个流程的。不过基本上,你能从工单里看到哪些活儿已经被处理过了,对吧,然后你可以从那里拿到这些代码,看看它用了哪些文档。另外还会有一段返修历史,是因为用错了零件导致的。所以你大致就能


[1:49:34]

get this view of okay well there's a bunch of documentation that we're maybe not hitting because we're getting all these comebacks from like using potentially the wrong part. Um so then there's that half of it and then there's some um documentation that we have um that's just not covered because we're using DTC codes that's just not covered inside of the warehouse. So we'll know that like certain documentation hasn't

得到这样一个视图:好吧,有一堆文档我们可能根本没覆盖到,因为我们看到这么多返修,很可能就是用错了零件造成的。这是其中一半。另一半是,我们手上有些文档根本没被覆盖到,因为我们用的那些 DTC 代码在仓库里压根没有对应内容。所以我们会知道,某些文档


[1:49:57]

really been hit at all. Um, so here it actually took a little bit more of a um, and it will do this sometimes. Here it used semantic expansion. Um, so it actually went in and did a much more comprehensive um, semantic expansion search. It's supposed to in this case use the outline. I think because it had the hierarchical URI um, it was probably able to do this a little bit better. Um but if I go back up um where did it show?

其实完全没被命中过。嗯,这里它实际上多做了一点……它有时候会这样。这里它用了语义扩展。所以它进去做了一次范围大得多的语义扩展检索。这种情况下它本来应该用 outline 的。我想是因为它有分层的 URI,所以做起来才稍微顺一些。不过如果我往上翻……它显示在哪来着?


[1:50:35]

Um yeah it well it was able to show here the field codes essentially um that it was missing. So um there was basically two field codes um that got a um yeah the headline mismatch or diagnosed with the wrong procedure. Um, so there's basically this mismatch um where a couple of these parts on the library document. Yeah. So there where the correct repair isn't um referred to basically. Um and that's causing a problem for the warehouse. Um

嗯,是的,它在这里能把缺失的那些 field code 列出来。所以基本上有两个 field code……嗯,是标题不匹配,或者说是用错误的流程做了诊断。所以基本上存在这种不匹配:库文档里的某几个零件……对,就是那里,正确的维修方式根本没有被提及。这就给仓库那边带来了问题。


[1:51:14]

and then there's other um DTC codes that occur in the field. Exactly. that have two um that have no code level documentation which are these two. So there's codes that are occurring which basically aren't documented that it was able to find. Um it's a shame here that it didn't use the outline it's supposed to do that because basically when it uses the outline it's able to traverse through and find all of the links a

然后还有另外一些在现场出现的 DTC 代码。没错。有两个……有两个是完全没有 code 级别文档的,就是这两个。所以有些代码在实际发生,但根本没有文档,这些它都能找出来。可惜的是这里它没有用 outline,它本来应该用的,因为它一旦用 outline,就能顺着遍历、把所有链接都找到,


[1:51:38]

little bit more efficiently. It wouldn't have done the comprehensive full text search crawl that took it a while in this case. Um, but even here because it had the hierarchical URIs and that semantic expansion on the um, full text search, it was able to eventually find it. Um, but it's a good lesson that when it does use the outline, it can come back faster because it can traverse out on the different links.

效率会高不少。那样它就不用做这次耗了不少时间的全文检索大范围爬取了。不过即便如此,因为它有分层 URI,加上全文检索上的语义扩展,它最终还是找到了。这是个很好的教训:当它真的用 outline 时,返回会更快,因为它可以沿着不同的链接往外遍历。


[1:52:02]

Um, the other thing here, the last one that I'll run because we only have eight minutes left. Um, I'll go ahead and copy it here. Um, and this one is primarily, if I go ahead and copy it, um, is going to leverage our, uh, theme shape. So, um, this is asking for common patterns across all our bulletins and recalls um,

嗯,另外一件事,这是我要跑的最后一个,因为我们只剩八分钟了。我把它复制过来。这一个主要是……我把它复制一下……主要会用到我们的 theme shape。所以这是在问我们所有技术公告和召回里的共性模式,


[1:52:26]

and how many of the cars sort of each affects. So, we have our work order history inside of our uh, warehouse. And uh basically what we want to find out is like which you know how does the sort of themes that we have correlate with um our work order history. Um and to do that it runs that theme pattern to be able to pull out the highle themes and then correlate it back to our work orders.

以及每一项大概影响了多少台车。我们的工单历史在仓库里。我们想搞清楚的基本上是:我们提取出的这些主题,跟工单历史之间是怎么关联的。为此它会跑 theme pattern,把高层次的主题抽出来,再关联回我们的工单。


[1:52:57]

And so it's a coverage question. So you can see it'll run the themes and then um eventually here it will go ahead and bring back um all of the all the relevant information. But basically what it's doing is it's it's finding the themes. It's going to go and pull sort of the types of fixes that it has in the work order history and then it's going to do a join of sorts to kind of group them under the right theme.

所以这是一个覆盖度的问题。你可以看到它会先跑主题,然后最终会把所有相关信息都带回来。但它做的事情基本上就是:先找出主题,再去工单历史里把各类维修方式拉出来,然后做一次类似 join 的操作,把它们归到正确的主题下面。


[1:53:31]

While this is running, are there any other questions about this theme shape or anything? I know other people there's Yes. Go ahead.

趁它在跑,关于这个 theme shape 或者别的还有什么问题吗?我知道还有其他人……有的,请讲。


[1:53:43]

Sorry. Say that one more time. how during the graph construction.

抱歉,能再说一遍吗?在 graph 构建过程中是怎么做的。


[1:53:56]

Yeah.

对。


[1:53:57]

Yeah. So the shapes that I that I are actually defined inside of our specs. So if you go here, our outline shape um our specs in our docs. So this was defined more thinking through like here's what we want it to look like for the agent. And then once we define that,

对。这些 shape 实际上是定义在我们的 spec 里的。你看这里,我们的 outline shape,我们文档里的 spec。定义的时候更多是在想:我们希望它对 agent 呈现出什么样子。定义好之后,


[1:54:17]

we come up with sort of the query structure that we want and we use that to inform our data model. Um, and all of the documents were loaded essentially um into that data model that I went over about an hour and a half before towards the beginning of the course. Um, it's that has containment tree with all the links between it.

我们就能得出想要的查询结构,再用它来指导我们的数据模型。所有文档基本上都是按那个数据模型加载进去的,就是我大概一个半小时前、课程开头讲过的那个。它有一棵包含关系树,以及之间的所有链接。


[1:54:37]

Um, and you can see here it'll um it'll bring back the um that it used themes. it didn't use search. Um, and then it uh also uh queried the um if I go up here, you can see the cars affected. So, it basically went to the warehouse with that connection shape um and it was able to sort of group them under um the different uh theme types. So, it's like that theme, the different theme types um that we have and then kind of with the

你可以看到这里它会带回来,说明它用了 themes,没有用 search。然后它还查询了……我往上翻一下,你能看到受影响的车辆数。所以它基本上是用那个 connection shape 去访问了仓库,然后把它们按不同的主题类型归了组。就是我们的那些不同 theme 类型,下面再挂上


[1:55:11]

cars and the work orders grouped underneath by percentages. Um yes, over here.

对应的车辆和工单,并按百分比分组。嗯,这边这位。


[1:55:21]

Yeah, either one. Do do I have a sense of accuracy of how well these do when the knowledge bases expand? Um, for this I haven't run any specific benchmarking on like exactly what I've shown you today. But I will say that when we do have customers that run these, they'll often come up with their own custom ontologies and then they will run benchmarks uh that will basically say like how effective is this query pattern against you know ordinary

对,哪个都行。我对知识库规模扩大之后这些方法的准确率有没有概念?就今天演示的这些内容,我没有跑过专门的 benchmark。但我可以说,我们确实有客户在用这套东西,他们往往会构建自己的定制 ontology,然后跑 benchmark,基本上就是评估这种查询模式相比普通的


[1:55:56]

vector search for example and you would use that to prove it out on on a specific type of data set. Here for this course this is more conceptual to understand kind of like the different shapes that you would use to help ground your data. Um, and it can always use work in terms of how you were to like build the skill, right? To make sure it guides through the right thing so it uses, you know, each step efficiently essentially.

vector search 效果如何,然后用它在某个特定类型的数据集上验证效果。今天这门课更偏概念,是为了理解你可以用哪些不同的 shape 来给数据做 grounding。而且在怎么构建 skill 这件事上永远有优化空间,对吧?要确保它能引导 agent 走对路径,让每一步都用得高效。


[1:56:23]

Are there any other questions?

还有别的问题吗?


[1:56:25]

Yep. Yeah.

好的。对。


[1:56:52]

Well, maybe I don't fully understand. So, you have um a connection point between between what? Between Yeah.

嗯,可能是我没完全理解。你是说你有一个连接点,在……在什么和什么之间?在……对。


[1:57:14]

Yeah.

对。


[1:57:17]

Yeah. No, we do not have so right now. So, now that I understand your question, the question is, is there a link between the connection structured data graph and the unstructured one? And there is not. In this case, it's they're completely unlin and the agent is sort of extracting, you know, a part from one or it's extracting like um yeah, basically like a code, right? And then referring back to the document. So, it's it's

对。不,我们没有,目前是没有的。现在我明白你的问题了,你问的是:结构化数据的 connection graph 和非结构化的那个之间有没有链接?答案是没有。在这个案例里,它们完全是没有关联的,agent 相当于是从一边抽取出……比如一个零件,或者说抽取出一个代码,对吧,然后再回去查文档。所以它是


[1:57:44]

using that to kind of do its own join in real time. You could create that linking and that could be valuable for more deterministic mapping. Um here we didn't do it and the main reason I didn't do it was really for speed of getting started because I did kind of want to provide code that would be easy and also model agnostic. Um and this is that and then if you wanted to later create those connections or those links you could.

用这种方式在实时做自己的 join。你完全可以把这种链接建起来,那对更确定性的映射会很有价值。这里我们没做,主要原因其实是为了上手快,因为我确实想提供一份容易跑起来、而且不依赖特定模型的代码。现在这份就是这样。如果你之后想建立那些连接或者链接,也完全可以。


[1:58:12]

Um any other questions? All righty. Well, thank you everyone. Hopefully that was uh informative. If you want to go back, the the workshop's available for you to take, it's it's online. I am going to have to take the anthropic key and the big query down eventually. I'll I'll leave the BigQuery one up uh for a while, but the anthropic one I'll have to eventually take down. So, unfortunately, you will need to provide your own key, but other

还有别的问题吗?好嘞。那么,谢谢大家。希望这次内容对大家有帮助。如果你想回头再看,这个 workshop 是开放的,你随时可以做,就在线上。不过 Anthropic 的 key 和 BigQuery 我最终得撤掉。BigQuery 那个我会多留一阵子,但 Anthropic 那个我最后必须撤掉。所以很遗憾,你需要自己提供 key,但除此之外,


[1:58:42]

than that, you should be able to take it just fine. Um, and I think that's all I have. So, I'll leave it for my next guest. Hopefully you guys have enough time to jump into your next session.

其他部分你都能顺利跑完。我想我要讲的就这些。我把场地留给下一位讲者。希望大家有足够的时间赶去下一场。


[1:59:06]

[music]

[音乐]