ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.61 · 全文

Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI

频道: AI Engineer
视频: https://www.youtube.com/watch?v=ZRM_TfEZcIo
原文语言: en
统计: 共 17 轮 · Paul 11 · Louis 6


[0:00] Paul

I spent 18 months turning my second brain into my living research memory. Let me explain. So, within my second brain, I currently have over 5,000 notes in Obsidian and another 5,000 notes in Readwise and some scattered in Notion and Google Drive. And all of this is growing on every week 250 files per month. And this is what I want. On the left, you can see my whole Obsidian vault, this huge mass. And whenever I start working on something such as an article, a new project, a new code base, a new feature, or whatever, I want to actually pull high-signal notes that are actually useful for my current work. And you would ask yourself, why not use directly Codex Cloud or Notebook LM? And I think it's that I am. But you need a system that sits between those harnesses and your second brain. Okay, so let's go back to the root of my problem, which is that I'm always losing my research. For example, my reading list is a graveyard. When I'm scrolling social media and I save that cool X post, a new article, a new new YouTube video, a GitHub repository, it doesn't matter. Whenever I actually want to start working on something, I never recall what I have in my second brain or I have to spend a ton of time actually finding meaningful notes that I can use in my work, right? And another problem that I have is that I want the system to actually be anchored into my personal notes, into my personal values, into my personal faith. I want the system to be personal to reflect my own thoughts, right? And that's why in today's video, Louis François and I will teach you how to build your own AI research OS. This also comes with code, so you can also try it out yourself. And I'm Paul Yushin. I'm the founder and CEO of Decoding AI, where I do a ton of content on courses on how to ship AI products. And I'm also the co-author of the LM Engineers Handbook bestseller. And the system, the AI research OS that I will teach you in this video is the system that I use in my daily work. And now I will pass the torch to Luis François.

我花了 18 个月,把自己的第二大脑变成了一个活的研究记忆系统。我来解释一下。在我的第二大脑里,现在 Obsidian 存了 5000 多条笔记,Readwise 里还有 5000 多条,另外还有一些零散地放在 Notion 和 Google Drive。这堆东西每周都在涨,每个月新增 250 个文件。我要的就是这个效果。左边你看到的是我整个 Obsidian vault,这么一大坨。每当我开始做点什么——写一篇文章、起一个新项目、搭一个新 code base、做一个新 feature,不管是啥——我都想真正地把那些对当前工作有用的高信号笔记捞出来。你可能会问,那为什么不直接用 Codex Cloud 或者 Notebook LM 呢?我确实在用啊。但你需要一个系统,夹在这些工具和你的第二大脑之间。好,那我们回到问题的根子上:我老是把自己的研究弄丢。比如我的阅读清单就是个坟场。我刷社交媒体的时候存了个很酷的 X 帖子、一篇新文章、一个新的 YouTube 视频、一个 GitHub repository,啥都有。可真到我要动手做事的时候,我从来想不起来第二大脑里有啥,或者得花一大堆时间才能找到能用在工作里的有意义的笔记,对吧?我还有另一个问题:我想让这个系统真正扎根在我个人的笔记里、个人的价值观里、个人的信念里。我想让这个系统是私人的,能反映我自己的想法。所以在今天这期视频里,Louis-François 和我会教你怎么搭建属于你自己的 AI 研究 OS。我们还附带了代码,你可以自己上手试。我是 Paul Iusztin,Decoding AI 的创始人兼 CEO,我做了大量关于如何打磨 AI 产品的内容和课程。我也是畅销书《LLM Engineer's Handbook》的合著者。这期视频里我要教你的这套系统——AI 研究 OS——就是我日常工作里真正在用的系统。现在我把话筒交给 Louis-François。


[2:05] Louis

Thanks, Paul. So, I'm Luis François Bouchard. I'm the co-founder and CTO of Towards AI, where we build educational courses. And I'm also the creator of What's AI, a YouTube channel where I explain AI engineering techniques. I used to explain AI research before. Now, I'm focusing on AI engineering. I'm also the author of the book Building AI systems for production. And before that, I was a PhD student. So, I honestly make research for a living. I used to do a PhD, as I said, in AI and doing tons of research and research work. Now, I build courses, I write videos, I research for videos, I build trainings for companies for a living. And all of these things that I do start with a very good research. And also leveraging tons of knowledge and insights that we get at Towards AI from building for clients. So, I have tons of notes as well, just like Paul. And we try to leverage them the best possible. And as you'll see, we'll build some sort of tool to leverage our second brain, where as you'll see, there will be some differences between how I use it and how Paul uses it. And that's the core goal of the repository that we built and on this project is that we want you to adapt it for your needs. The whole goal is how can we make research better, but more specifically, how can we better leverage what we have? So, let's dive into it. And first, we need to figure out which tool to use and when, because this whole research system that we built is not for every query. If you just need a fast answer, like a few quick questions or just something where that that you would just Google, basically. Well, obviously, just Google it or ask ChatGPT, Cloud, whichever system you want. But, the problem when doing that is that if you have a lot of following up question or it's a bigger project that you need to build on and have basically a very long context or tons of information to share, relying on ChatGPT isn't ideal. And it also means that you are fully dependent on the architecture that OpenAI or ChatGPT's team built.

谢谢,Paul。我是 Louis-François Bouchard,Towards AI 的联合创始人兼 CTO,我们做教育课程。我还是 YouTube 频道 What's AI 的主理人,在那儿讲 AI engineering 的技术。我以前讲 AI research,现在更聚焦在 AI engineering 上。我也是《Building AI Systems for Production》这本书的作者。在那之前我是个博士生。所以说实话,搞研究就是我的本职。我前面说了,我读过 AI 方向的博士,做了一大堆研究和研究工作。现在我做课程、写视频、为视频做调研、给公司做培训,这就是我的饭碗。而我做的所有这些事,起点都是一份非常扎实的研究。同时也得借力我们在 Towards AI 给客户做项目时积累的大量知识和洞察。所以我跟 Paul 一样,也有海量笔记,我们都想把它们用到极致。等下你会看到,我们要搭一个工具来调用我们的第二大脑,而且我用它的方式和 Paul 用它的方式会有些不同。这正是我们搭的这个仓库、这个项目的核心目标——我们希望你能按自己的需求去改它。整件事的目标是:怎么把研究做得更好?更具体地说,怎么把我们手头已有的东西利用得更好?那我们就开始吧。首先得搞清楚什么时候用哪个工具,因为我们搭的这整套研究系统并不是为每个 query 准备的。如果你只是要个快速答案,比如几个简单问题,或者那种你随手 Google 一下就行的事——那显然直接 Google 就好,或者问 ChatGPT、Claude,随你用哪个系统。但这么干的问题在于,如果你有很多追问,或者这是个更大的、需要持续往上搭的项目,需要很长的 context、要喂一大堆信息进去,那靠 ChatGPT 就不太理想了。而且这也意味着你完全被 OpenAI 或者 ChatGPT 团队搭的那套架构绑死了。


[4:20] Louis

So, the next step here is to ask yourself for a more complex problem, do you need to act quickly or do you want to build some next feature and and do something very difficult? If you just have a small repo for a quick change or write one article, just do one thing that you know won't be repeatable that much, definitely use Codex or Cloud Code or some agent that you trust. Sometimes, you need to keep on digging to make it better, to improve efficiency, optimize it more. And so, typically, when you have to do that, you want your research sources, your research to stick and to be able to refer to them in the future. So, if you want a process like this where the sources that you find, the notes that you take stick around in time and have an agent be able to leverage that efficiently and being able to come back to these information, to ask follow-up questions, to digest content even more. And right now, for instance, when I make a new video, I want also the agent and the system to understand the previous videos I made to not duplicate content, to not repeat myself, and to refer to some other content. In this case, there are some tools that are very interesting that you might have tried before, like NotebookLM, that is super powerful to do research, to digest content efficiently, and to come back to it. But, the problem with NotebookLM is that is Well, first, the main problem is that you don't own it. You cannot do anything you want with it. You cannot personalize as much as possible. It's not agent native. And it's obviously weak for coding tasks since it's just browser based. So, it's far from ideal from something that Paul and I needed and that most AI engineers need in general. So, if you need your agents to be able to leverage all you do, uh whether it is a big research, a new video, whatever you write, you do, you code, you typically want your other agents, your other projects to be able to leverage what you learn from what you just did. And one thing that we advise especially for production, obviously for product, is to build some sort of retrieval rag pipeline with vector databases. But, this needs an infrastructure. It's not really human-friendly to be able to digest quickly, to check notes, to make edits.

所以下一步,对于更复杂的问题,你要问自己:你是想快速搞定,还是想搭点下一个 feature、做点很难的东西?如果你只是有个小 repo 要改一处地方,或者写一篇文章,就做一件你知道不太会重复的事,那绝对该用 Codex、Claude Code,或者某个你信得过的 agent。但有时候你得不断往深里挖,把它做得更好、提升效率、做更多优化。这种情况下,你通常希望你的研究来源、你的研究能沉淀下来,将来还能回头去引用。所以如果你想要那种流程——你找到的来源、你做的笔记能随时间留存,让一个 agent 能高效地调用它们、能回头看这些信息、能追问、能把内容嚼得更透——比如现在我做一期新视频时,我也想让 agent 和系统理解我之前做过的视频,这样就不会做重复内容、不会自我重复,还能引用一些别的内容。这种场景下,有些工具就很有意思,你可能之前试过,比如 NotebookLM,它做研究、高效消化内容、回头复看都特别强。但 NotebookLM 的问题在于……嗯,首先最大的问题是它不归你所有。你没法拿它随心所欲,没法尽可能地个性化。它不是 agent native 的,而且因为它就是个浏览器里的东西,做编码任务明显很弱。所以离 Paul 和我需要的东西、离大多数 AI engineer 普遍需要的东西,差得还很远。所以如果你需要你的 agent 能调用你做的一切——不管是一份大研究、一期新视频,还是你写的任何东西、你做的、你 code 的——你通常希望你的其他 agent、其他项目也能从你刚做完的事里借力。我们尤其建议在生产环境里、显然也是为产品考虑时,去搭一套带 vector database 的 retrieval rag 流水线。但这需要一套基础设施。它对人不太友好,没法让你快速消化、查笔记、做修改。


[6:40] Louis

It's hard to inspect by hand. You need to build everything around it. It's definitely far from ideal for just something I want to use on a daily basis. Obviously, it's super powerful at scale, very interesting especially in a product. But, as I said, this project is for us. And I don't want something live, super professional as a product. I just want something I will use and that my agents and different projects can leverage as best as possible. So, the last question to ask ourselves here is that if you want everything there but more personalization, so a personalized research assistant that builds some sort of Wikipedia that compounds over time and and is easily inspectable and usable where you have tons of sources, documents, videos, comparisons, implementations, new research, new topics that you keep on adding and that you keep on wanting to leverage and review easily, this is where you may want to build something yourself. And in our case, we built a personalized research OS that we will share in this talk with exactly what we built and how. But, the downside is that it definitely needs a bit more setup than just opening Cloud Code. Right now, the main problem with using Cloud Code and other agentic tool is that you give Codex links, PDFs, and different information, for example, my most recent Loop Engineering video, and then the next session you use Codex, you have to paste it all again or ask it to use skills. And whatever structure that Codex or ChatGPT, whatever tool that you use, build on the fly to leverage what you did, the scripts it ran, the scripts it had, you all lost it or kept it inside a skill that you have to ask it to reuse, and that usually isn't ideal and just grows and grows over time. And the problem is that all this information that you give to the model is not the bottleneck. The bottleneck is how can you leverage it in the future? Meaning that with an agent, the context window becomes everything, the database, the file system, the memory, the reasoning space. It has to do it all, and when you stop the conversation, it loses everything. And the thing is that we don't need necessarily to provide more and more and more context for a better research. You need a proper memory and context management, and ideally some personality with it, especially in my case when I do videos.

它很难靠手去检查。你得围着它把一切都搭起来。对于我只想日常随手用的东西来说,它绝对远谈不上理想。当然,它规模化之后超级强,尤其在产品里非常有意思。但我前面说了,这个项目是给我们自己用的。我不想要一个上线的、像产品那样超级专业的东西。我只想要一个我自己会用、而且我的 agent 和各种项目都能尽量去借力的东西。所以这里要问自己的最后一个问题是:如果你想要上面这些全都有、但要更个性化——也就是一个个性化的研究助理,它会搭出某种随时间不断累积复利的 Wikipedia,而且容易检查、容易用,里面有海量的来源、文档、视频、对比、实现、新研究、新主题,你不停往里加,而且不停想去调用、想方便地复看——那这就是你可能想自己动手搭点东西的时候了。我们这边搭的是一个个性化的研究 OS,这次分享里我们会把我们到底搭了什么、怎么搭的都告诉你。但缺点是,它确实比直接打开 Claude Code 要多一点搭建工作。现在用 Claude Code 和其他 agent 工具的主要问题是:你给 Codex 喂链接、PDF、各种信息,比如我最近那期 LLM Engineering 视频,然后下次你再用 Codex,你得把这些全重新粘一遍,或者让它去用 skills。而不管 Codex 还是 ChatGPT、不管你用哪个工具,临时搭出来用以调用你做过的事的那套结构、它跑过的脚本、它用过的脚本,你要么全丢了,要么塞进了某个 skill 里、还得每次专门让它去复用——而这通常都不理想,还会随时间越滚越大。问题在于,你喂给模型的所有这些信息并不是瓶颈,瓶颈是:你将来怎么调用它?也就是说,对一个 agent 来说,context window 成了一切——数据库、文件系统、记忆、推理空间,它全得一肩挑,而你一结束对话,它就把一切都丢了。事情的关键是:我们并不一定要喂越来越多的 context 才能换来更好的研究。你需要的是一套像样的记忆和 context 管理,最好再带点个性,尤其对我做视频的情况来说。


[9:16] Louis

So, what we did is that we decided to build a system with plain files, mostly markdown files, that we can leverage easily and that agents can leverage easily. I won't detail it very much here because Paul will talk about it in depth. And as I said, Paul has like 5,000 or something notes. I have just a few hundred, but that's just to say that we need to consider that we didn't start from nothing. We already both had some sort of large database. In my case, I made hundreds of videos and I take many notes. So, I still need to leverage these years of content that I already made and tons of meetings that I have with my team, with clients when we build for them that I want to leverage as well because we learn a lot by building for people. We have highlights from interesting tweets that I see, interesting articles that I see, and I want all my projects to be able to leverage my agent skills. So, I decided to pivot and instead of having a folder for Cloud Code skills and having all my meeting recaps in Granola and having years of notes on Apple Notes and the tweets on the saved Chrome tab, instead, I moved everything automatically into Obsidian. It's just a note reader, obviously, so you don't have to use that. You can just save it locally, but I used Codex to set up everything so that Granola is automatically saved there, my notes are now on Obsidian just because it's a nice UI, I like it, and I can use it from my phone, my computer, my Windows, Mac, everything. So, anyways, I moved everything to Obsidian, which means that it's saved locally in my file system, which means it's basically my companion for researching and building everything I build now. And what we built, obviously, leverages that. We built a repo called AI Research OS for this workshop where it's basically just skills for Cloud Code and Codex with plugins to be able to do a very deep research about a topic or a simpler search or distillation, different tools that you can use. The most useful and complete one will be the research tool that I use, for example, when I kick off a new video topic. And the goal of this repo is to have you implement it, install the cloud plugins from it, and tune it to your needs. Right now, it can connect to, as I said, Obsidian with my local notes. It can use Readwise, Notebook LM, your GitHub repos, any links that you send for GitHub or YouTube videos, and web links, obviously, and documents. But there are tons of things missing, as we will discuss in the end, that you can easily implement, just asking Cloud Coder or Codex to do so. Like for example, I implemented the YouTube video transcript in honestly a few seconds, just one prompt.

所以我们做的是,决定用纯文件——主要是 markdown 文件——来搭一套系统,我们自己能轻松调用,agent 也能轻松调用。这里我不细讲了,因为 Paul 待会儿会深入说。前面提到,Paul 有 5000 条左右的笔记,我只有几百条,但说这个只是想强调:我们得考虑到,我们俩都不是从零开始的,我们都已经有了某种规模可观的数据库。我这边呢,做了几百期视频,也记了很多笔记。所以我还是得去调用这些年我已经做出来的内容,还有我跟团队、跟客户(给他们做项目时)开的大量会,这些我也想用上,因为给别人做东西的过程中我们学到了很多。我们还有从看到的有意思的推文、有意思的文章里摘下来的 highlights,我想让我所有的项目都能调用我的 agent skills。所以我决定来个转向:与其搞一个文件夹放 Claude Code skills、把所有会议纪要放 Granola、把这些年的笔记放 Apple Notes、把推文放 Chrome 收藏标签页里,我反过来把所有东西都自动挪进了 Obsidian。它显然就是个笔记阅读器,所以你不一定非得用它,你存本地就行,但我用 Codex 把一切都配好了——让 Granola 自动存进去,我的笔记现在都在 Obsidian 上,纯粹因为它 UI 漂亮、我喜欢,而且我手机、电脑、Windows、Mac,哪儿都能用。总之,我把一切都挪进了 Obsidian,这意味着它本地存在我的文件系统里,意味着它基本上就成了我现在做研究、搭东西的随身伴侣。我们搭的东西显然就借力于此。我们搭了一个叫 AI Research OS 的 repo,专为这次 workshop——它本质上就是给 Claude Code 和 Codex 用的一堆 skills,配上插件,让你能对一个主题做非常深的研究,或者做更简单的搜索、做提炼,各种你能用的工具。其中最有用、最完整的是那个研究工具,比如我开一个新视频选题时就会用它。这个 repo 的目标就是让你把它实现出来、装上它的 cloud 插件、再按你的需求调。它现在能连——我前面说了——连我本地笔记所在的 Obsidian,能用 Readwise、Notebook LM、你的 GitHub repos、你发来的任何 GitHub 或 YouTube 视频链接、当然还有网页链接和文档。但还有一大堆缺的东西,我们最后会讨论,那些你都能轻松实现,直接让 Claude Code 或 Codex 去做就行。比如那个 YouTube 视频转写,说实话我几秒就实现了,就一个 prompt。


[12:07] Louis

It's super easy for Codex to implement it. So, the thing is that this whole repository and this whole project is a very useful companion for my own work. But, as I said, it implements tons of features and state-of-the-art context management and memory management techniques that I believe AI engineers need to know. And now, Paul will dive into all this three-layer system that we built with the raw content, the index that I mentioned, and the wiki-like synthesized version of all your notes, all your research, all your work. So, he'll cover everything we did, how it ended up, and show how to use it.

对 Codex 来说实现起来超简单。所以呢,这整个仓库、这整个项目,对我自己的工作来说是个非常有用的伴侣。但前面也说了,它实现了一大堆功能,以及我认为 AI engineer 需要了解的最前沿的 context 管理和记忆管理技术。接下来 Paul 会深入讲我们搭的这套三层系统——原始内容、我提到的那个 index,还有把你所有笔记、所有研究、所有工作综合起来的那个类 wiki 的版本。所以他会把我们做的一切、最后变成了什么样、以及怎么用都讲一遍。


[12:48] Paul

Okay. So, now I want to go over the three versions of our system and how it progressed over time, and most importantly, why we added more complexity. So, in the first version, we wanted to scope it just to create lessons for our agent engineering course. So, we wanted to keep it super simple, where we had as input a topic and a research MD as output. So, within the input, we had the topic plus a set of golden links, which were manually handpicked by us. We applied this deep research algorithm, and we had as output a static research MD file. And if we go zoom into the architecture, we first scraped the links of the golden links, right? Because we already know them and we use them as seed for context for the deep research algorithm, which was a really powerful technique because we had more context on how to frame our questions. And during the query rounds, we basically used the very classic deep research algorithm where we had one main agent, the orchestrator, which created multiple questions based on the initial topic and the scraped context. And each agent managed its own question and used Gemini grounded in Google to to query basically Google and gather multiple resources and each agent gather these resources, which returned multiple links and created some executive summaries of each link. And then [snorts] it passed all this information back to the agent to the main agent where the main agent basically aggregated all this information into a summarized way so it did not exploded the context. And we did this for three rounds. So basically after three rounds of generating six queries per round, we ended up with like 40-50 links in total. So you can imagine that there's a lot of noise over there. So that's why we also applied a ranking algorithm where we wanted to like find the the information with the highest signal. And basically we compared each source against the topic, the initial topic of the user. And like that, we fully scraped only the top K elements based on the ranking score. And for the rest of the of the links, we just kept the summaries. And then we compiled everything into this research MD file as a single flat file which we used for each lesson of our course in in our particular use case.

好。那现在我想过一遍我们这套系统的三个版本,看看它是怎么一步步演进的,最重要的是,我们为什么要往里加更多复杂度。在第一个版本里,我们想把它的范围限定在只为我们的 agent engineering 课程生成课时。所以我们想保持超级简单:输入是一个主题,输出是一份 research MD。在输入端,我们有主题,加上一组黄金链接(golden links),这些是我们手工精挑出来的。我们对它跑这套深度研究算法,输出是一份静态的 research MD 文件。如果放大去看这套架构:我们首先抓取那些黄金链接的内容,对吧?因为我们本来就知道它们,就把它们当作深度研究算法的种子 context,这是个非常强的手法,因为我们对怎么去框定问题就有了更多 context。在 query 轮次里,我们基本上用的是非常经典的深度研究算法:有一个主 agent,也就是 orchestrator,它基于初始主题和抓来的 context 生成多个问题。每个 agent 负责自己那个问题,用 grounded in Google 的 Gemini 去查 Google、搜集多份资料,每个 agent 搜集这些资料、返回多个链接,并为每个链接生成执行摘要(executive summary)。然后它把所有这些信息传回给主 agent,主 agent 基本上就是把这些信息聚合成一份摘要的形式,这样就不会把 context 撑爆。我们这么跑了三轮。所以基本上,每轮生成六个 query、跑三轮之后,我们总共拿到大概 40 到 50 个链接。你可以想象,这里头噪声相当多。所以我们又加了一套排序算法,想找出信号最高的那些信息。基本做法是把每个来源和那个主题——用户的初始主题——做对比,这样我们就只对排名分数最高的 top K 个元素做完整抓取,剩下的链接就只留摘要。然后我们把一切编进这份 research MD 文件,作为一个扁平的单文件,在我们这个具体用例里给课程的每一课时用。


[15:27] Paul

But as you can imagine, it was pretty limited. For the course, it worked great right away. We generated 35 lessons really quick, but we wanted more. So, we started to aim this deep research loop to the second brain as well, right? Before, it was targeting only the public web, which made this a pretty generic, and we had to manually find all those golden links. So, by aiming this deep research loop to the second brain, where we basically organically keep track of all the information that of all the research that we really want and is filtered by us, we can organically gather all those golden rings into our deep research algorithm. So, let's look at how this new algorithm looks like. It's basically the same loop, right? But now we target our own sources instead of just the public web. Now, for input, we have only the topic because we don't need the golden links. As I said, the golden links are actually a reflection of our second brain system. In theory, you can also add them if you really want to, but that's the beauty of this new strategy because you can just put as input some topic and we'll find everything that it needs. And then we use this topic only as seed for for for the context to generate the the queries. And now we do the same deep research algorithm, right? The same query rounds, but instead of targeting only the public web, now we plugged in all our second brains, such as the our Obsidian, our Readwise, our Notebook LM, our GitHub. You can also use, for example, Gemini deep research for this, like similar to how we we use notebook LM or you can extend this with whatever you want. For example, YouTube, uh Google Drive, Notion, or whatever makes sense on your infrastructure. The idea is that now we are target our queries from the deep research algorithm to our second brain plus the public web. And after we apply the same algorithm such as ranking, fully scraping, summaries, and compile everything into this research MD file. But now we have another problem, right?

但你也能想到,它相当受限。对课程来说,它一上来就跑得很好,我们很快就生成了 35 节课时,但我们想要更多。所以我们开始把这套深度研究循环也对准第二大脑。在那之前,它只面向公开网络,这就让它相当通用,而且我们还得手工去找所有那些黄金链接。把这套深度研究循环对准第二大脑——我们在那儿有机地记录着所有我们真正想要、并经我们筛选过的信息和研究——这样我们就能有机地把所有那些黄金链接汇集进我们的深度研究算法。那我们来看看这套新算法长啥样。它基本上是同一个循环,对吧?但现在我们对准的是自己的来源,而不只是公开网络。现在输入端我们只有主题,因为不需要黄金链接了。我前面说过,黄金链接其实就是我们第二大脑系统的一种映照。理论上你真想加也可以加,但这正是这套新策略的妙处——你只要把某个主题作为输入,它就会把需要的一切都找出来。然后我们只把这个主题当作生成 query 的 context 种子。接着我们做同样的深度研究算法,对吧?同样的 query 轮次,但不再只对准公开网络,现在我们把所有第二大脑都接了进来,比如我们的 Obsidian、Readwise、Notebook LM、GitHub。你也可以用,比如 Gemini deep research 来做这个,跟我们用 Notebook LM 的方式类似,或者你想接什么都能扩。比如 YouTube、Google Drive、Notion,或者任何在你的基础设施上说得通的东西。思路就是,现在我们把深度研究算法发出的 query 对准我们的第二大脑加上公开网络。之后我们再套同样的算法——排序、完整抓取、摘要,把一切编进这份 research MD 文件。但现在我们又有了另一个问题,对吧?


[17:48] Paul

This this research MD file is static. It's a pile of static data. And usually research is not static, right? So, after you end up with with this file, you most often realize that you want to ask another question or some information is stale and you don't need it anymore. Or or basically, you want more out of this research MD file. And we which means that you need to start all of this from scratch. And the operation that I showed you uh above is an extensive operation. It consumes a lot of tokens and it takes a lot of time. So, you don't want to run it from scratch. And that's why you need to add a wiki layer on top of it. And that's why V3 of this system is actually a deep research algorithm plus an LM knowledge base on top of it, aka the wiki layer. So, the new algorithm looks like this. So, we have sources in and a wiki out. And the sources, as I said before, can be like Obsidian, notebook LM, Google Drive, or YouTube, Notion, even custom URLs, right? That that's also powerful as well, where you use tools such as a Bright Data to to parse basically any single page application, any type of site, any type of public information that's out there. We can put it in. And then you apply the same deep research algorithm. You store everything into raw files, right? Instead of compiling everything into research and D file, now we store each file individually. And we create an index out of all these files. And ultimately, we generate a wiki on top of it, and which we can query. We can query basically the wiki plus the index. Okay, so this is just the high-level architecture of the new system. Let's zoom into it. So, what do I want to start with is that you should forget the infra structure you think you need, such as vector databases, knowledge graphs, semantic search, text search. All that is beautiful, but add a lot of complexity, especially for like this personal wikis, personal research operating systems that you want to use very lightly. So, I want a system just based on files, right? A simple mechanism that's that's very rooted into how your computer works.

这份 research MD 文件是静态的,它就是一堆静态数据。可研究通常不是静态的,对吧?所以你拿到这个文件之后,多半会发现你想再问个别的问题,或者某些信息已经过时了、你不再需要了,又或者你就是想从这份 research MD 文件里榨出更多东西。这意味着你得从头把这一切再来一遍。而我上面给你看的那个操作是个很重的操作,它消耗大量 token,也很费时间。所以你不会想从头跑一遍。这就是为什么你需要在它上面加一层 wiki。所以这套系统的 V3 其实就是深度研究算法,再加上盖在它之上的一个 LLM 知识库,也就是那层 wiki。新算法长这样:来源进、wiki 出。来源嘛,我前面说了,可以是 Obsidian、Notebook LM、Google Drive,或者 YouTube、Notion,甚至是自定义 URL,对吧?这一点也很强——你可以用 Bright Data 这类工具去解析基本上任何单页应用、任何类型的站点、任何外面公开的信息,都能塞进来。然后你套同样的深度研究算法,把一切存成原始文件,对吧?不再把一切编进一份 research MD 文件,现在我们把每个文件单独存。我们再从所有这些文件生成一个 index。最终,我们在它上面生成一个 wiki,并且可以去查询它——我们基本上可以查询 wiki 加上 index。好,这就是新系统的高层架构。我们放大来看。我想先说的是:你该忘掉那些你以为自己需要的基础设施,比如 vector database、knowledge graph、semantic search、text search。这些都很美,但会带来大量复杂度,尤其对这种你想很轻量地用的个人 wiki、个人研究操作系统来说。所以我想要一套就基于文件的系统,对吧?一种简单的机制,深深扎根于你的电脑本来就是怎么运作的。


[20:06] Paul

And that's why we'll create all the system just based on files and just based on references. So, no database, just a simple index based on references. And how how this works? We have an agent that reads an index.yaml file that's basically a catalog of your all your data plus the summaries of of each source and some metadata around it. For example, here on the right, you can see part of an index.yaml file that contains 10 sources and 38 wiki pages as derivatives of these sources, where uh we can see there into the sources list of the YAML file the first uh source, for example. And as you can see, it has like the the the link to the original file plus some metadata, such the origin, the title, the authors, the the the publication date, the summary, and and things things like this, which can be flexible, right? And the next step is that based on this index.yaml file, we need to point to all the wiki pages, to all the wiki derivatives, to all the raw sources. So, basically, this index.yaml file is an entry point for our agent, right? It's what we will give to our agent to actually reason on how to find our data. It's an index, ultimately, right? So, the next step is to understand how the wiki actually looks like. So, on the left, you can see the high-level structure of the wiki, where we have the raw folder, the wiki folder, and the index. In the raw folder, we actually just have the raw data, which is immutable. You don't want to touch that. The And the index points to everything that we need. And in the wiki, we actually have derivatives created by the LLM, which contains things such as comparisons between multiple concepts, entities, or just simple notes as a reflection of our questions or repositories that we ingested, and we can create multiple notes based on a on on a repository, right? Or open questions that based on our questions that LLM couldn't answer yet. And everything that you can analyze on top of your raw data. And on the right, based on Obsidian, we can see like the sub graph reflected just based on in this index. And this is just like the first iteration, but as the the the the wiki grows, you can see connections made between entities and concepts. For example, concepts are things such as tool registry, context compaction, sandboxing, or entities are open code, closed code, MCP, right? So, as you can see, you can beautifully can start visually and practically create connections.

所以我们整套系统都只建在文件和引用之上。不用数据库,就是一个简单的、基于引用的 index。那它具体怎么运作呢?我们有一个 agent,它会去读一个 index.yaml 文件——这个文件本质上就是你所有数据的目录,加上每个来源的摘要和一些围绕它的元数据。比如右边你能看到一段 index.yaml,里面包含 10 个来源,以及由这些来源派生出的 38 个 wiki 页面。在 YAML 的 sources 列表里,我们能看到第一个来源长什么样:它有指向原始文件的链接,加上一些元数据,比如出处、标题、作者、发表日期、摘要等等,而且这些字段是灵活可变的,对吧?下一步是,基于这个 index.yaml,我们要让它指向所有的 wiki 页面、所有的 wiki 派生物、所有的原始来源。所以本质上,这个 index.yaml 就是 agent 的入口;我们就是把它喂给 agent,让它据此去推理该怎么找到我们的数据。说到底,它就是个 index,对吧?再下一步是搞清楚 wiki 本身长什么样。左边你能看到 wiki 的高层结构:我们有 raw 文件夹、wiki 文件夹和那个 index。raw 文件夹里放的就是原始数据,它是不可变的,你不想去碰它。index 则指向我们需要的一切。而在 wiki 里,我们放的是由 LLM 生成的派生内容,比如多个概念或实体之间的对比,或者只是一些简单的笔记——是对我们提的问题、对我们摄入的仓库的一种反映。我们可以基于一个仓库生成多条笔记,对吧?或者是一些 open questions——那些基于我们提问、但 LLM 暂时还答不上来的开放问题。总之,凡是你能在原始数据之上做的分析,都放在这。右边是在 Obsidian 里、基于这个 index 反映出来的子图。这只是第一轮迭代,但随着 wiki 不断生长,你能看到实体与概念之间建立起的连接。比如概念有 tool registry、context compaction、sandboxing,实体有 open code、closed code、MCP 等等。所以你看,你能很漂亮地、既直观又务实地把这些连接建起来。


[22:49] Paul

Now, the next obvious question is how do we actually query this wiki? So, as I said, the agent will have as input this index.yaml file, which contains summaries and metadata about each source. But what happens next, right? The next step is actually to look into the source wiki page. Where the source wiki page is like an executive summary of each page. Which is basically not just a summary, but a more expanded summary of of each source. And sometimes the agent just looks into this, gets what it needs, and goes back. Which is very token efficient, right? And if it doesn't find within this uh source wiki page, we also need links into the wiki derivative, such as concepts, entities, notes, comparisons, and so and so forth. And only if it doesn't find the necessary information up to this point, it needs and it actually reads the whole raw source, right? Which basically contains the whole article, the whole paper, the the the whole video, or whatever. And this makes just through pure referencing and creating this simple hierarchy, this makes everything very token efficient. Now, the beautiful part is that this wiki is actually alive, right? For example, every question leaves a trace into your wiki. So for every question, the LLM can create a new concept file, a new notes [snorts] file, a new comparison file. And every question is is tracked into a log. So basically, the the the wiki doesn't evolve only when you ingest new data or do a deep research round, it actually evolves as you start talking with it, right? That's the beautiful part, actually. And like that, you can see a true reflection of of yourself, of what you haven't understood, of all your questions from the past. And the beautiful part is that the the wiki is never frozen, right? Similar to the research entity files. At any point, you can ingest a new custom link that you think that you need into the wiki, or even run a new deep research round. Or as I said previously, the wiki keeps evolving just purely based on your questions.

那么接下来一个很自然的问题是:我们到底怎么查询这个 wiki?前面说了,agent 的输入是这个 index.yaml,里面有每个来源的摘要和元数据。但之后会发生什么呢?下一步其实是去看 source wiki page。这个 source wiki page 相当于每个页面的执行摘要——它不只是一个简短摘要,而是对每个来源更展开一些的摘要。很多时候 agent 只看这一层,拿到它要的东西,就回去了,这非常省 token,对吧?如果它在 source wiki page 里没找到,我们还需要指向 wiki 派生物的链接,比如概念、实体、笔记、对比等等。只有到这一步还没找到必要信息时,它才会真的去读整份原始来源——也就是整篇文章、整篇论文、整段视频之类的完整内容。靠纯粹的引用、加上这套简单的层级结构,整件事就变得非常省 token。现在最妙的一点是,这个 wiki 其实是活的,对吧?比如说,每一个问题都会在你的 wiki 里留下痕迹。对每一个问题,LLM 都可以新建一个概念文件、一个笔记文件、一个对比文件,而且每个问题都会被记进日志。所以这个 wiki 并不是只在你摄入新数据、或者跑一轮 deep research 时才进化——它会随着你开始和它对话而不断进化,对吧?这才是真正妙的地方。这样一来,你就能看到一面真实的自我镜像:你哪些东西还没搞懂、你过去提过的所有问题,都映照在里面。而且最棒的是,这个 wiki 永远不会被冻结,跟前面说的 research 实体文件一样。任何时候,你都可以把一个你觉得需要的自定义链接摄入 wiki,甚至再跑一轮新的 deep research。或者像我前面说的,wiki 会单纯靠你的提问持续进化下去。


[25:01] Paul

And another important thing to understand is that this wiki doesn't sit on top your of your entire second brain, right? For example, in my particular use case, I use the PARA method coined by Tiago Forte where all my data is structured between project, areas, resources, and archive. Where all my notes resources that I save, sources that I save yet article or whatever are just piped directly to the resource a flat list. And whenever I need something, I just references them into projects and areas, right? And like this, Obsidian is just an immutable snapshot that a LLM never touches, right? So, this is my data. I don't really want the LLM to touch my personal notes that I manually write, right? So, then how can we actually put this wiki to use, right? So, as I said, we have the big Obsidian snapshot which is our global second brain. And then whenever we start to work on a new project, we reference this second brain through this deep research algorithm that I explained, and we scope it down to our own project, right? So, basically, whenever we want to start working on something, we run this deep research loop or we start ingesting some particular repositories, articles, notes, and so on and so forth. And we usually do that through a set of skills plugged into a harness. And a project can be basically anything such as writing a new article, doing a new video, doing a set of slides. I I applied this technique doing this slide. Or you can even apply it for something more complex such as writing a book, doing a course, or or keeping track of a whole code base, right? You can also you use it for that. So, basically, a project can be anything where you want, as I said initially, to transform research into work. The project is the work, and your second brain is the research. So, now I want to show you a few demos. So, what you need to do is go to the AI research OS workshop repository, and here you can find all the skills required to run what we presented into this presentation.

还有一点很重要,要理解的是:这个 wiki 并不是凌驾在你整个第二大脑之上的,对吧?拿我自己的用法来说,我用的是 Tiago Forte 提出的 PARA 方法,我所有的数据都被组织成 project(项目)、area(领域)、resource(资源)、archive(归档)四类。我保存的所有笔记、资源、来源——无论是文章还是别的什么——都直接灌进 resource,是一个扁平的列表。当我需要什么的时候,再把它们引用进 project 和 area 里,对吧?这样一来,Obsidian 就是一个不可变的快照,LLM 永远不去碰它。这是我的数据,我真的不希望 LLM 去动我亲手写的个人笔记,对吧?那我们到底怎么把这个 wiki 用起来呢?前面说了,我们有那个庞大的 Obsidian 快照,它是我们全局的第二大脑。然后每当我们开始做一个新项目,就通过我刚讲的那套 deep research 算法去引用这个第二大脑,把它收窄到我们自己这个项目的范围里,对吧?所以本质上,每当我们想动手做点什么,就跑这个 deep research 循环,或者开始摄入某些特定的仓库、文章、笔记等等。我们通常是通过一组插在 harness 上的 skill 来完成这件事的。而一个项目可以是几乎任何东西,比如写一篇新文章、做一期新视频、做一套幻灯片——我做这套幻灯片就用了这个技巧。你甚至可以把它用在更复杂的事情上,比如写一本书、做一门课、或者追踪一整个代码库,对吧?这些也都能用它来搞。所以本质上,项目可以是任何「你想把研究转化成产出」的场景。项目就是产出,而你的第二大脑就是研究。好,现在我想给你们看几个 demo。你要做的就是打开那个 AI research OS workshop 仓库,在里面你能找到运行我们这次演示所需的全部 skill。


[27:17] Paul

And everything is packed as a cloud code plugin, but you can very easily tweak it and install it with any other harness. And also in the read me, you can find details on how to install all all the other dependencies, because the thing is that this system is dependent on tools such as Obsidian, Readwise notebook elements, and so on and so forth. So, you need to set up specific CLIs or authentication issues. But, I don't really want to waste any of your time with setup issues, and I want to go straight directly into the examples. So, I prepared here three examples. The first one is a research on one of my previous articles on agentic AI engineering, and within these files, I have a brand dump of everything that I knew I wanted to talk on this subject. And on top of that, I also added a few references that I knew 100% that I want to add into this wiki. And what we need to do to actually trigger the the algorithm on top of these files is to open up a cloud session, and then just call the skill, the research skill, and point it it to this file. And that's it. Everything else is baked directly into the skill. It will understand my intent that I want to create a wiki on agentic harness engineering just looking at these files and looking at the topic, and it will know that before starting the deep research algorithm, it actually needs to scrape this information to to use it as context when it frames the questions for the deep research algorithm. Now, we need to wait a bit for for the agent to reason on top of it. And I will actually just put it on auto mode to speed speed this up, right? I use this hundreds of times, so I know it won't delete anything from my computer or it won't do anything weird. Okay, so now this is the most important part, right? So, it asked me how deep I want the deep research algorithm to be. We have light, deep, fast. This mostly controls how many questions you want to run per one round and how many rounds you want to run. And usually light or fast is more than enough because remember this process consumes a lot of tokens. So, you kind of need to do the deep one only when you you really want to look over tons and tons of notes. And for this use case, I will just pick the light one to to to speed this up. And in this use case, it just does one round of three queries, right? And for the fast one, it does two rounds of three queries. So, I would just keep it keep it around that spectrum.

所有东西都打包成了一个 Claude Code 插件,但你可以非常轻松地改一改,把它装到任何别的 harness 上。另外在 README 里,你能找到如何安装其它所有依赖的细节,因为这套系统是依赖一些工具的,比如 Obsidian、Readwise、Notebook LM 之类。所以你需要配置一些特定的 CLI 或者处理认证问题。不过我真的不想拿配置问题浪费你们的时间,我想直接进入示例。我这里准备了三个示例。第一个是针对我之前一篇关于 agentic AI engineering 文章做的研究。在这些文件里,我做了一次脑暴倾泻,把我知道自己想就这个主题谈的所有东西都写下来了。在此之上,我还加了几个我百分百确定要放进这个 wiki 的参考资料。要在这些文件上触发这套算法,我们要做的就是开一个 Claude 会话,然后调用那个 research skill,把它指向这个文件,就这样。其它一切都已经烘焙进 skill 里了。它会理解我的意图——我想就 agentic harness engineering 建一个 wiki,它光看这些文件、看这个主题就懂了;而且它知道,在启动 deep research 算法之前,它得先把这些信息抓取下来,当作给 deep research 算法构造问题时的 context。现在我们得稍等一下,让 agent 在这上面推理。我干脆把它切到 auto 模式来加快速度,对吧?我用过这玩意几百次了,所以我知道它不会删我电脑里的东西,也不会干什么奇怪的事。好,现在到了最重要的一步,对吧?它问我希望 deep research 算法挖多深。我们有 light、deep、fast 几档。这主要控制你每一轮想跑多少个问题、想跑多少轮。通常 light 或 fast 就绰绰有余了,因为记住,这个过程会吃掉大量 token。所以只有当你真的要在海量笔记上翻找时,才需要用 deep 那档。这个用例里,我就选 light 来加快速度。这一档就跑一轮、三个查询,对吧。而 fast 那档是跑两轮、每轮三个查询。所以我建议你就在这个区间里调。


[29:52] Paul

And now the process will will take around 10 to 20 minutes to actually look around my Obsidian, to look around my Readwise, my Notebook LM, and run those queries on top of this. And I actually run this, right? And now let's open this wiki that we created based on the prompt before in Obsidian and let's look what's inside the the wiki. So, we have three big objects, the raw files, which are basically a raw copy of what we found, the index, which contains all the references to towards the wiki, right? We can also in Obsidian have this beautiful a subgraph where we can very quickly understand what is going on. This is created purely based on this this file. And then the most interesting part is inside the wiki. Where we have comparisons, concepts, entities, and sources. The sources contain the executive summary of of our raw sources. So, the LLM doesn't really need to every time when it reads them, and it needs them to to read the raw sources, but it needs just the executive summaries, which are computed just one time during the ingestion. And for example, for the comparisons, it understood out of the box that it needs to do comparisons between like adjective rag versus file systems or compaction versus recursive language models or or or anything of interest based on on the sources. And the most interesting part is actually the concepts. So, it automatically extracted all the concepts that we need to understand from this pile of resources. For example, if we open the agent loop resources, we automatically can look and get this beautiful summary containing like graphs, tags, and explaining us everything that we need on this topic. And we can do the same on all the concepts from from from this wiki or all the entities from the wiki and so and so forth. Okay, so now let's go to the second example. It It's a simple example where I want to learn more on harness engineering and how harnesses work. So, in in this prompt over here, I just want to ingest the three open source repositories on open code, Pi, and Hermes.

现在这个过程大概要花 10 到 20 分钟,它会真的去翻我的 Obsidian、翻我的 Readwise、翻我的 Notebook LM,然后在这些之上跑那些查询。我其实已经事先跑过一遍了,对吧。现在我们在 Obsidian 里打开这个根据前面那条 prompt 创建出来的 wiki,看看里面有什么。我们有三大块:raw 文件,本质上就是我们找到的内容的原始副本;index,里面装着所有指向 wiki 的引用;在 Obsidian 里我们还能看到这张漂亮的子图,让我们能很快搞清楚整体在发生什么——这张图纯粹是基于这个文件生成的。而最有意思的部分在 wiki 里面。里面有 comparisons(对比)、concepts(概念)、entities(实体)和 sources(来源)。sources 里装的是我们原始来源的执行摘要。所以 LLM 并不需要每次都去读原始来源——它要用的时候,只需要这些执行摘要就够了,而这些摘要是在摄入时一次性算好的。再比如 comparisons,它开箱即用地就明白自己该做哪些对比,比如 agentic RAG 对比 file systems、compaction 对比 recursive language models,或者基于这些来源的任何有意思的对比。而最有意思的其实是 concepts。它自动把我们需要理解的所有概念都抽取了出来。比如我们打开 agent loop 这个资源,就能自动看到这份漂亮的摘要,里面有图、有标签,把这个主题上我们需要的一切都讲清楚了。这个 wiki 里所有的概念、所有的实体,我们都能这么干,等等等等。好,现在来看第二个示例。这是个很简单的例子,我想多了解一些 harness engineering、以及 harness 是怎么运作的。在这条 prompt 里,我只想摄入三个开源仓库——open code、Pi 和 Hermes。


[32:16] Paul

And I don't want to do deep research at all, right? I just want to ingest those repositories and explore topics such as the general architecture, agents architecture, sub agents, memory system, and the agent permission flow. And let's run the research on top of this prompt. And now what the research will do will clone automatically all these repositories and will explore all the repositories on the topics that I gave here and it will create notes at the individual level of each repository on how they work, on how how the architecture works at the repository level and then we can create higher level notes, right? Within the wiki derivatives and compare all the architectures or create aggregate architectures on on like the general trends of all those harnesses and basically explore and learn everything that we want about harness engineering directly from the code. And again, usually I just do auto mode and let it do its own gist, but I already run this, right? So here is the wiki for these GitHub repositories and again, we have the raw files and we have the index and the wiki. And here probably within the repos we can see all the three repositories and for example, inside the open core repositories we can see empty files explaining everything that we need on particular topics such as the permission flow, the memory system and so on and so forth and we can see this in all the other repositories and the most interesting part here is actually that we have comparisons on all of these, right? So we can actually understand the differences within the architecture within these harnesses or we also have all the concepts extracted from these repositories and we can understand what are the key architectural decisions from these repositories. So we can go crazy with this and this is super useful if you want to, for example, write your own harness. And the third examples is the simplest one in in reality, which is just based on ingesting some some simple links, right? And again, I will just exit the the previous run and I want to run this example from scratch.

而且我完全不想做 deep research,对吧?我只想摄入这几个仓库,去探索一些主题,比如总体架构、agents 架构、sub agents、记忆系统,还有 agent 的权限流。我们就在这条 prompt 上跑这次研究。接下来这次研究会自动克隆所有这些仓库,然后按我这里给的主题去逐个探索每个仓库,并在每个仓库各自的层面上生成笔记——讲它们是怎么工作的、在仓库这一层架构是怎么运作的;然后我们可以在 wiki 派生物里生成更高层的笔记,对吧,把所有架构拿来对比,或者就所有这些 harness 的总体趋势做出聚合式的架构,基本上就是直接从代码里把我们想了解的关于 harness engineering 的一切都探索并学到手。还是那句话,我通常就开 auto 模式,让它自己跑,但我已经事先跑过这个了,对吧。这里就是这几个 GitHub 仓库的 wiki,同样地,我们有 raw 文件、有 index、有 wiki。这里在 repos 下面我们能看到全部三个仓库,比如在 open code 仓库里,我们能看到一些笔记文件,把我们需要的一切都讲清楚了,比如权限流、记忆系统等等;其它仓库里也都能看到这些。而这里最有意思的部分,其实是我们对这些全都做了对比,对吧。所以我们真的能搞清楚这几个 harness 在架构上的差异,我们也有从这些仓库里抽取出来的所有概念,能搞明白这些仓库背后的关键架构决策是什么。所以这玩意你可以玩出花来,如果你想——比如说——自己写一个 harness,它就超级有用。第三个示例其实是最简单的,无非就是摄入几个简单的链接,对吧。同样地,我会先退出上一次的运行,然后从头跑这个示例。


[34:41] Paul

And here, I just want to show you that you can use this also with a very basic setup where I just want to ingest three custom random links. And as before, I just pass this prompt and it will start the research process. And I want to highlight that you can run example two ending on the GitHub repositories and this example without any other setup like without setting up Obsidian, Readwise, or anything else. You can just install this plugin and run this examples because it here is not dependent on any other service than Git and using curl to get this this URLs. And again, here if we go into Obsidian and explore the wiki created out of this example three, we we can see the index, we can see the wiki itself with all the sources, right? One executive summary for for each source and all the concepts, extracted entities, and so on and so forth. And the idea is that as you start asking questions on top of this, everything starts to get more interesting. So now, let's assume that we want to ask for example a question on harness engineering based on the wiki created out of the GitHub repositories. So what we have to do is just again hit the research field pointed to the wiki that we just created, and then just ask our question. And let's assume that I want to learn more on sandboxing, more exactly how remote sandboxing works and how is plugged into the heart. This can be basically any any any other question. And now what it will happen, it will basically query this wiki, it will give you an answer, and based on this, you can also start creating notes, comparisons, or maybe it will extract and find new entities that you care about. So basically, it will start updating the wiki.

这里我想给你们演示的是,你也可以在一个非常基础的配置下用它——我这里只想摄入三个随便挑的自定义链接。跟之前一样,我把这条 prompt 一传,它就开始研究流程。我想强调一下:第二个示例(停在 GitHub 仓库那个)和这个示例,你不需要任何额外配置就能跑——不用配 Obsidian、Readwise,或者别的任何东西。你只要装上这个插件就能跑这些示例,因为它们除了 Git、以及用 curl 去拉这些 URL 之外,不依赖任何其它服务。同样地,如果我们进 Obsidian 去看这个示例三创建出来的 wiki,就能看到 index,能看到 wiki 本身,里面有所有 sources——每个来源一份执行摘要——还有所有概念、抽取出来的实体等等。关键在于,一旦你开始在这之上提问,一切才开始变得更有意思。那现在,假设我们想基于刚才用 GitHub 仓库建出来的那个 wiki,提一个关于 harness engineering 的问题。我们要做的,同样是再次调用那个 research 字段、把它指向我们刚建好的 wiki,然后直接问我们的问题。假设我想多了解一些 sandboxing,更具体地说,是远程 sandboxing 是怎么工作的、它是怎么接进 harness 的。这其实可以是任何别的问题。接下来会发生什么呢——它会去查询这个 wiki,给你一个答案,并在此基础上,让你可以开始创建笔记、做对比,或者它会抽取并发现一些你在意的新实体。所以本质上,它会开始更新这个 wiki。


[36:49] Louis

All right. So now, where is this project going? What is it? What should you do? First, it's still rough at some points. Like it needs more connectors, as you saw. We need to add Google Drive, Notion, Slack, and tons of other connectors that could be useful to you. But to be honest, it's not really useful to me and my current workflow or to Paul. So we didn't add them yet, because the core of this project is to be useful for us and for you to take over and add whatever you need. And the other main goal of this project is to teach memory and context management. So all these extra features aren't really useful towards that purpose. There are other some weaknesses, like it's hard to know which sources are outdated or weak or strong compared to some other system that we built. So we know we can improve this, but again, it's not really the priority here. And lastly, it's still obviously a builder workflow. You use it through cloud code and Codex, and it's just to me in the terminal and I just really like it. So it's not a final polished product with a nice UI, nice UX. And honestly, that's by design. So we don't really care about this, because our goal is to teach AI engineering. It's not to build the next best product. Still, we have a few next improvements we want to do very shortly, from having a stronger linting to a better memory compaction, because that's a big issue and it's just very complicated in general to manage memory correctly, and the state of the art is always progressing there. We, as I said, need better source provenance to trust the sources and be able to rank them properly and reuse them properly if needed and be able to know access quickly as a user if this source is relevant or not. And we have other next improvements to do. But those are mostly for optimization and for the future. And the thing is that we actually build all of that into another product that you can even build yourself. Because we created a course called Agent Engineering where we build a similar deep research system with a writing and research agent where we build a system with the same goal to be able to learn best AI engineering practices. It's a very in-depth course where I assume it takes around 60 hours to complete with a final project being the multi-agent system I just described that you'll build for yourself. So if this presentation and the demo repo that you saw was interesting, please consider checking out the Towards AI Academy with our courses on there including the Agent Engineering course to learn more on the best practices when building around and with agents.

好。那么这个项目接下来要往哪走?它到底是什么?你又该做点什么?首先,它在某些地方还很粗糙。比如像你们看到的,它需要更多 connector。我们得加上 Google Drive、Notion、Slack,以及一大堆别的、可能对你有用的 connector。但说实话,这些对我自己当前的工作流、或者对 Paul 来说,并没有那么有用,所以我们暂时还没加——因为这个项目的核心,是对我们、以及对接手并按需扩展的你们有用。这个项目另一个主要目标,是讲清楚 memory 和 context 管理。所以那些额外功能对这个目的其实没什么帮助。它还有一些别的弱点,比如很难判断哪些来源是过时的、薄弱的还是可靠的——比起我们做过的另一套系统而言。我们知道这块可以改进,但同样,这在这里并不是优先级。最后,它显然还是一套面向构建者的工作流。你是通过 Claude Code 和 Codex 来用它的,对我来说就是在终端里用,而我就是很喜欢这样。所以它不是一个有漂亮 UI、漂亮 UX 的、最终打磨好的产品。老实说,这是刻意为之的。我们其实并不在意这个,因为我们的目标是教 AI engineering,不是去打造下一个最好的产品。尽管如此,我们还是有几项很快就想做的改进,从更强的 linting,到更好的 memory compaction——因为那是个大问题,而把 memory 管对,整体上本来就非常复杂,而且这块的最前沿技术一直在往前推进。还有像我说的,我们需要更好的来源溯源,这样才能信任这些来源、能恰当地给它们排序、在需要时恰当地复用它们,并且作为用户能快速判断某个来源到底相不相关。我们还有别的一些后续改进要做,但那些大多是为了优化、为了将来。而关键是,我们其实已经把这一切都做进了另一个产品里,那个产品你甚至可以自己动手搭出来。因为我们做了一门叫 Agent Engineering 的课,在课里我们搭了一套类似的 deep research 系统,带一个写作 agent 和一个研究 agent,目标也一样——让你能学到最好的 AI engineering 实践。这是一门非常深入的课,我估计大概要 60 小时才能学完,期末项目就是我刚才描述的那套多 agent 系统,你将亲手为自己搭出来。所以如果这次演讲、以及你们看到的这个 demo 仓库让你觉得有意思,欢迎去看看 Towards AI Academy 上我们的课程,包括这门 Agent Engineering 课,去深入学习围绕 agent、以及用 agent 来构建时的最佳实践。