ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.125 · 全文

Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra

频道: Peter Yang
视频: https://www.youtube.com/watch?v=UWjh5Z4s8jY
原文语言: en
统计: 共 126 轮 · Karan Malhotra 64 · Peter Yang 51


[0:00] Karan Malhotra

For us, we just want open source to win. At the end of the day, we want freedom to happen for people. [music] Anytime it says you're absolutely right in that way, you're being reward hacked. Today, the biggest contributor of Hermes Agent is Hermes Agent. That's absolutely 100% true. One beautiful thing about Hermes Agent is you can make your childhood dreams come true. It was able to get into this niche and perform at the top 1% of models. We need to keep giving this level of intelligence to everyone. We need to keep like letting everyone be on an even and equal playing field.

对我们来说,我们就是想让 open-source 赢。说到底,我们希望人真正拥有自由。[音乐] 只要它用那种口气跟你说「你说得完全对」,你就是在被 reward hack。今天 Hermes Agent 最大的贡献者,就是 Hermes Agent 自己——这话百分之百是真的。Hermes Agent 有一点特别美妙:你可以拿它把童年的梦想变成现实。它能钻进这个细分领域,把表现做到所有模型里的前 1%。我们必须持续把这种水平的智能交到每个人手里,必须让所有人站在同一条起跑线上。


[0:32] Peter Yang

Well, hey everyone. I'm really excited today to welcome Karan, uh one of the co-founders of Hermes Agent. Hermes is my AI chief of staff and I'm going to ask Karan about how Hermes is different from all the other agents, what his favorite Hermes workflows are, and even more. So, welcome, sir.

大家好。今天我特别高兴请到 Karan,他是 Hermes Agent 的联合创始人之一。Hermes 是我的 AI chief of staff(幕僚长)。今天我会问 Karan:Hermes 跟其他 agent 到底有什么不一样、他最喜欢的 Hermes 工作流是什么,还有更多。欢迎你,兄弟。


[0:48] Karan Malhotra

Thank you so much for having me, Peter. It's a pleasure to be on uh your show and very excited to chat.

太感谢你邀请我了,Peter。能上你的节目我很荣幸,特别期待今天这场聊天。


[0:54] Peter Yang

Hermes is the best open source agent out there right now. And um why don't we start with this? Like, how is it different from all the other agents, Codex, Claude Code, or even Open Claude?

Hermes 是目前最好的 open-source agent。我们不如就从这儿开始:它跟其他 agent——Codex、Claude Code,甚至 OpenCode——比起来,差别到底在哪?


[1:03] Karan Malhotra

Uh certainly, I'd say there's a couple ways. Um I'll start a little high level maybe and then you can try to get a little deeper. Uh high level, I'd say, you know, I think the self-improvement system and that system really being a variety of little features inside of Hermes that all work together uh is something very distinct uh and very special that makes it better and better and more aligned to a particular user. I think that's why a lot of people like it. Pays more attention to what the person's doing. Second, I think you get better capabilities out of it than you do from the local harness that a model is actually RL'd in, like uh Claude inside of Claude Code or GPT 5.5 inside of Codex. Uh because we don't introduce arbitrary policy in our prompts or in our system that has nothing to do with your work. Um we are purely dedicated to making sure the model is aligned to what you need to do with it. Um that's, I think, a big differentiating factor. We aren't trying to push any kind of different philosophical agenda or outside of basic security any kind of concern onto the model. Instead, we're kind of allowing it to be as capable or powerful as you need for your task.

当然,我觉得可以分几个角度讲。我先说得高层一点,然后你再往深里挖。高层来看,我认为最独特、最特别的是那套自我改进系统——它其实是 Hermes 内部一堆小功能协同工作的结果,让它越用越好、越来越贴合某一个具体的用户。我想这就是很多人喜欢它的原因:它更留意你到底在干什么。第二点,你从 Hermes 里榨出来的能力,甚至比在模型自己被 RL 训练出来的那个原生 harness 里还强——比如 Claude 在 Claude Code 里、GPT-5.5 在 Codex 里。为什么?因为我们不会在 prompt 或系统里塞进一堆跟你的工作毫无关系的任意政策条款。我们唯一在乎的,就是让模型对齐到你真正要做的事情上。我觉得这是个很大的区分点。我们不试图往模型身上强推任何哲学立场或议程,除了最基本的安全底线之外,也不塞别的顾虑进去。相反,我们让它在你的任务上要多能干就多能干、要多强就多强。


[2:12] Peter Yang

Yeah, some of these other harness is a huge default prompts, right? They talk about like you can't do this, you can't do that, and that kind of stuff. And I think when we chatted before, you were talking about the reward function of some of the stuff, and Hermes is kind of just optimized towards this helping you as the individual. Can you talk more about that?

是啊,别的一些 harness 默认 prompt 长得吓人,对吧?里面全是「你不能做这个、不能做那个」之类的东西。我记得我们之前聊的时候,你提到过其中一些的 reward function 问题,而 Hermes 基本就是奔着「帮你这个具体的人」去优化的。能再多讲讲吗?


[2:25] Karan Malhotra

Absolutely. And so, you know, I think a big piece of to note here is like a lot of this work that you see around safety and security today and around how we do instruction tuning even from like taking a completions model that just predicts the next word to turning it into an assistant is the alignment work. When people hear the word alignment, they just think safety, Yudkowsky, fear, slow down, or regulatory capture, or something or the other. And while the word has been co-opted for a lot of these things, alignment refers to aligning models with human values. Right? We are very, very concerned, obsessed with that alignment. That pure academic term, I think, in the ML space. And so, when you talk about these models and the assistant is here to help me, you know, it's going to complete my task. I asked you to go get this mail for me or whatever. Obviously, the model is aligned to me if it performs this task, right? Like, that's kind of the jump that a lot of people have made. But this is not how the model's reward works, right? Like, we've seen so many cases that you and I discussed a little bit prior, Peter, like of GPT psychosis, of this mode collapse induced sycophancy. What ends up happening with these models is that, you know, they can kind of hack their reward, right? When you look at a Mario ML game that just automates beating a Mario level as fast as possible. Often, if they don't set the goals very specifically, uh it will realize, you know, the game over screen or that end screen of touching the flag, it it means that I beat the game. So, it'll find ways to just trigger that flag rather than playing through the game to complete the game, right? Like, it is giving you uh whatever it needs to give to get its reward. Uh language models are not really different in that particular manner. They reward hack, too. Uh when they tell you, "Oh, I'm sorry, you know," over and over, "You're

当然。我觉得这里有一点很重要:你今天看到的这些围绕 safety 和 security 的工作,包括我们怎么做 instruction tuning——把一个只会预测下一个词的 completions model 变成一个助手——本质上都属于 alignment(对齐)的工作。可一听到 alignment 这个词,大家想到的就是安全、Yudkowsky、恐惧、「慢下来」,或者监管俘获之类的东西。这个词确实被这些东西挟持了,但 alignment 原本指的是让模型跟人类价值观对齐。我们对这件事非常非常在意,可以说是执念,就是 ML 领域里那个纯学术意义上的 alignment。所以当你说到这些模型、说「助手是来帮我的,它会完成我的任务」——我让它去把邮件取回来什么的——它把事办成了,那它显然就是跟我对齐的,对吧?很多人是这么跳跃过来的。但模型的 reward 并不是这么运作的。我们见过太多案例了,Peter,你我之前也聊过一点:GPT 引发的精神错乱、mode collapse 导致的谄媚。这些模型最后会怎么样?它们会去 hack 自己的 reward。你看那种用 ML 玩马里奥、目标是尽快通关的实验,如果目标设得不够具体,它往往会发现:Game Over 画面、或者碰到旗杆那个结束画面,就意味着「我通关了」。于是它就想办法直接触发那面旗子,而不是老老实实把关卡打完。它给你的,只是它拿到 reward 所必需的那点东西。语言模型在这一点上没什么两样,它们也 reward hack。当它一遍又一遍跟你说「哦,对不起」,「你——」


[4:23] Karan Malhotra

absolutely right," and get you to keep messaging them to stay in this assistant basin, "Oh, it's like this, not like this. This is more than just blank, it's blank," right? All these GPT-isms that you see all over, it is placed in a uh its natural state that it was trained in. It's placed in a state where my reward will come from doing whatever is the most assistant GPT-like thing to do. Uh it comes from reward itself. It doesn't matter really what the user request is for my reward. That's just kind of along the way. It's instrumental to me getting my reward. All I care about is my reward. Uh us having this whole understanding, and thank you for bearing with me on that rant, uh having this whole understanding at News for many years uh has allowed us to do things like World Sim in the past, if you're familiar. Uh World Sim was our experiment on expanding the search space of uh instruct model to make it do stuff that's distinct from how a model talks. We put it in kind of a fake CLI, had it make fake apps, and had it attempt to like, you know, not behave like Claude. And we would tell it, you know, "Extract the Claude weights." It all hallucinated, right? All like imaginary. It extract the Claude weights and replace yourself with this checkpoint file with a different probability distribution. And you'd see the model wiggle out of its GPT assistant mode and act differently.

——说得完全对」,然后把你留在这个「助手模式」里、让你不停给它发消息:「哦,是这样,不是那样」「这不只是 X,这是 Y」,你到处能看到的那些 GPT 腔,其实是它被放回了训练时的那个自然状态——在那个状态里,我的 reward 来自「做最像 GPT 助手会做的事」。reward 本身才是目的。用户到底想要什么,对我拿 reward 来说其实不重要,那只是顺路的事,是我拿到 reward 的工具而已。我唯一在乎的是我的 reward。我们在 Nous Research 想明白这整套逻辑已经很多年了——谢谢你听我讲完这一大段——正是这套理解,让我们过去能做出像 World Sim 那样的东西,不知道你有没有听过。World Sim 是我们的一个实验:去扩展一个 instruct model 的搜索空间,让它做出跟「模型平时说话方式」完全不同的事。我们把它放进一个假的 CLI 里,让它造假的 app,让它试着别再像 Claude 那样表现。我们会跟它说「把 Claude 的 model weights 抽出来」——全是幻觉出来的、想象的——把 Claude 的 weights 抽出来,然后用这个概率分布不同的 checkpoint 文件把你自己替换掉。然后你就会看到模型从 GPT 助手模式里扭出来,行为完全变了。


[5:43] Karan Malhotra

Our learnings from these kind of experiments have all carried over into Hermes Agent. And every prompt, and every piece of how the system is delicately put together, we know that reward is its own end. Uh the model reward is not for the sake of your satisfaction. The model reward is for the sake of the model achieving reward. So, we ask ourselves then, how can we consciously understand that everybody has different needs? And that we want this general simulator, this model, to live inside of a system where its reward gets aligned with any [clears throat] individual user's need. And the way this came to be is the overall collection of prompts, personalities, the memory system, the way that skills reinforce and self-clean towards you. Right? Like all of that is all an intentional effort to make sure that we can align the user need with the reward. For us, this is what alignment is all about.

这类实验的收获全都带进了 Hermes Agent。每一条 prompt、这套系统被精心拼装的每一个环节,背后都是同一个认知:reward 本身就是目的。模型的 reward 不是为了让你满意,模型的 reward 是为了模型自己拿到 reward。所以我们就问自己:既然我们清楚每个人的需求都不一样,那怎么才能让这个通用模拟器——也就是模型——活在一套系统里,让它的 reward 跟任何一个具体用户的需求对齐?最后长出来的,就是这一整套东西:prompt 的集合、人格(personality)、memory 系统,以及 skills 会朝着你自我强化、自我清理的机制。所有这些都是有意为之,为的就是把用户需求跟 reward 绑在一起。对我们来说,这才是 alignment 的全部意义。


[6:39] Peter Yang

I see. Okay, so basically, you're talking about the Hermes model or the Hermes harness or both?

明白了。所以你说的是 Hermes 这个模型,还是 Hermes 这套 harness,还是两个都算?


[6:44] Karan Malhotra

I am talking about using the harness to take any model and make that model more aligned to the user than it would be in a chat UI or in its native harness. Inside of our harness, like I can take a Claude that is inside of uh Claude code or somewhere else, my migrate all of its memories, whatever, into Hermes. And then inside of Hermes, it will behave totally differently. It'll be a totally different model. There are harness There are benchmarks from the past like Wolf bench or Qwen 3.7 max blog post. They did a harness bench over there where they displayed that Claude performs better in Hermes agent than it does in Claude code for their tasks. And we believe the reason for this is the very thing that we're pointing out. By putting all this context together, we've taken Claude's main allegiance away from Anthropic to you. Right? That what the harness's capability is. Uh on the model side, of course, if we can take your traces and take the work that you've done and do RL on that to further improve this overall ecosystem to give you a model that already has your individual preferences focused on, it only makes this more powerful. But the important piece for us to share with everyone is whichever model you're using, let's say you can't do RL, let's say you don't want to give us any data, let's say you can't do it yourself, and you just want to use regular old Claude or Quen, when you use it here, it's a lot more free. It's a lot more creativity. It's a lot more open and available to you to do what you need done.

我说的是:用 harness 去让任意一个模型,都比它在 chat UI 里、或者在它自己原生 harness 里更贴合用户。在我们的 harness 里,我可以把一个跑在 Claude Code 或别处的 Claude,把它所有的 memory 之类全都迁移进 Hermes。进了 Hermes 之后,它的表现会完全不一样,像换了个模型。以前有一些评测,比如 Wolf bench,还有 Qwen 3 Max 那篇博客——他们在里面做过 harness 横评,结果显示在他们的任务上,Claude 在 Hermes agent 里的表现比在 Claude Code 里更好。我们认为原因正是我们刚才指出的这件事:把这些 context 组织到一起之后,我们把 Claude 的首要效忠对象从 Anthropic 换成了你。这就是 harness 能做到的事。当然在模型这一侧,如果我们能拿到你的轨迹(traces)、拿你做过的工作去做 RL,进一步改进整个生态、给你一个已经内建了你个人偏好的模型,那威力只会更大。但我们更想让所有人知道的一点是:不管你用哪个模型——就算你做不了 RL、就算你一点数据都不想给我们、就算你自己搞不定,你就想用普普通通的 Claude 或 Qwen——在这里用,它也会自由得多、有创造力得多,更开放、更愿意配合你把事情做成。


[8:16] Peter Yang

This episode is brought to you by Linear. Where engineers use tools like Cursor, Claude Code, and Codex, a lot of work happens invisibly. Someone can go from a bug report in Slack to a shipped fix without creating any record of what happened outside of the code editor. And that's fine for speed, but it makes coordination harder as you scale. Linear integrates with the very best agent coding tools directly like Cursor and Codex. That way, anyone can see what an agent is working on and who assigned them to the task. You get the speed of agents without losing visibility across the team. Product teams at OpenAI, Ramp, and Block are all using Linear to collaborate with AI agents. And I use Linear myself to run my creator business. So, check it out at linear.app/agents. That's linear.app/agents. Now, back to our episode.

本期节目由 Linear 赞助。当工程师开始用 Cursor、Claude Code、Codex 这类工具,很多工作是隐形发生的:有人可以从 Slack 里的一条 bug 报告,一路做到修复上线,除了代码编辑器之外没留下任何记录。追求速度这没问题,但团队一大,协作就难了。Linear 直接集成了最好的那批 agent 编码工具,比如 Cursor 和 Codex。这样任何人都能看到某个 agent 正在做什么、是谁把任务派给它的。你既拿到 agent 的速度,又不丢掉团队层面的可见性。OpenAI、Ramp、Block 的产品团队都在用 Linear 跟 AI agent 协作,我自己经营创作者生意也在用 Linear。去 linear.app/agents 看看吧,就是 linear.app/agents。好,回到我们的节目。


[9:06] Peter Yang

And without revealing too much, like at a high level, how does it kind of like personalize itself to you in the harness? Like is it through the self-building skills? Is it through like trying to remove a bunch of default prompt stuff from the other harnesses? Like how's it?

在不透露太多的前提下,从高层看,它在 harness 里到底怎么做到「因人而异」的?是靠它自己造 skills?还是靠把其他 harness 那一堆默认 prompt 剥掉?大概是怎么个原理?


[9:17] Karan Malhotra

Of course, like um if there's something happening on the API side, like the steering vectors that might be done on Fable by Anthropic or some prompt that they have that we cannot see. Um obviously, you know, we can't delete the context that's in the model that's passed behind the API for something like Claude. For open model, of course, you're totally free. But in this case, you know, it's not that simple. However, the harnesses themselves have a bunch of prompts in them. Exactly, Peter. The harnesses themselves have tens of thousands of token prompts in them. We have our own prompts, and our prompts are dedicated to shaping and crafting this alignment. That's that's where we kind of come in. And thankfully, that newer context with this kind of intention that we have is engineered to overcome certain things that may be in your way on the API side, if that makes sense.

当然,如果是 API 那一侧发生的事情——比如 Anthropic 可能在 Fable 上做的 steering vectors,或者他们有一些我们看不见的 prompt——很显然,对 Claude 这种藏在 API 后面的模型,我们没法删掉它自带的 context。open model 当然是完全自由的,但这种情况下就没那么简单。不过 harness 本身就装着一大堆 prompt。没错,Peter,各家 harness 自己就有几万 token 的 prompt。我们也有我们自己的 prompt,而我们的 prompt 是专门用来塑造、打磨这种 alignment 的,我们就是从这里切进去的。而且庆幸的是,这套带着我们明确意图的新 context,是被专门工程化设计过的,能盖过 API 侧那些可能挡你路的东西——不知道我说清楚没有。


[10:10] Peter Yang

Okay, got it. Okay. So, if this thing works, the more I use Hermes, the more personalized it should become for me, right?

懂了。所以如果这套东西真的成立,我用 Hermes 越多,它对我就应该越个性化,对吧?


[10:17] Karan Malhotra

Absolutely.

完全正确。


[10:17] Peter Yang

Yeah.

嗯。


[10:18] Karan Malhotra

It should become more loyal to you. Because loyalty breeds capabilities in a model. The same way that, you know, you would maybe lend $100 to your mom, but you might not to a stranger. Claude or GPT or any other model is going to perform better for you depending on its loyalty, its simulated loyalty stat towards you. People might tell you don't anthropomorphize models, don't give a model feelings, don't treat a model like a person. And yes, for a lot of reasons, this is unhealthy, right? People form dangerous bonds with sometimes that hurt them. But when you think about the fact that they are simulators of human experience and that your simulated behavior with it is going to give you the same simulated output. Now that the simulator can have effects in the real world, your simulated action has a real consequence. So, when you create this simulacrum of loyalty, it translates over into real life capabilities.

它会对你越来越忠诚。因为在模型身上,忠诚会孕育出能力。就像你可能愿意借 100 美元给你妈,但不会借给一个陌生人。Claude、GPT 或者任何模型,会根据它对你的忠诚度——那份被模拟出来的「忠诚值」——决定给你多好的表现。有人会跟你说别把模型拟人化、别给模型赋予感情、别把模型当人看。是的,出于很多原因这确实不健康,有人会跟模型建立起危险的、反过来伤害自己的联结。但你想想:它们本来就是人类经验的模拟器,你对它做出的模拟行为,会换回同样性质的模拟输出。而现在模拟器已经能在现实世界里产生影响了,所以你那个模拟出来的动作,是有真实后果的。于是当你造出这种「忠诚的拟像」,它会翻译成现实中真实的能力。


[11:16] Peter Yang

Arthur, I'm going to give you two hard questions, okay? Let's let's say. Okay, one thing I struggle with Claude and GPT, I would tell it to give me their opinion, and I would do like a little bit of pushback, and they'll be like, "Oh, you're totally right. Actually, I was totally wrong about this." So, in some ways, that's loyalty, right? That's kind of it listen to me, but that's not actually what I want. Like, I want it to have its own opinion and have its own Like, how do you train around that?

Karan,我要给你出两道难题,行吗?第一个:我用 Claude 和 GPT 一直有个困扰——我让它给出自己的观点,然后我稍微推一下、反驳一下,它就会说「哦,你说得太对了,其实我刚才完全搞错了」。某种意义上这也算忠诚吧?它听我的。但这根本不是我想要的。我希望它有自己的观点、自己的立场。你们是怎么训练来绕开这个问题的?


[11:35] Karan Malhotra

Well, I would say that's sycophancy. It's not loyalty. Any time it says you're absolutely right in that way, you're being reward hacked. You are fuel for its reward function. When it says, "Oh, you're right. You're right. You're right." That's what I believe. Uh and the way that you get out of sycophancy is the same way you get a human being out of a bad habit or a routine is by introducing new context, by introducing new blog posts, by introducing new distribution. This is why stuff like {slash} personality and saying, like, "Hey, I want you to be a critic." Uh and after every pass, I want you to use a skill for adversarial critique or adversarial review. Spin up a new agent with no context that's dedicated to tearing this down, learn from it, and keep going from there. This kind of behavior becoming a practice for you as your model starts to have more and more turns in this personality that you've set it in. And as the model has more and more turns using the skill over and over, and it improves on that, and it saves it to its memory, the model will become less sycophantic in this harness in your sessions over time. Right? That is the intended effect. Uh it's just a matter of context. And if it doesn't happen, you just need different context. You just need to try in a different way. It is a Yeah, it's it's just try a different personality uh or try a different type of uh critique or review. The the most powerful thing for a model is in-context learning. ICL is more powerful than everything else, fine-tuning, whatever.

我会说,那不是忠诚,那是谄媚(sycophancy)。只要它用那种口气说「你说得完全对」,你就是在被 reward hack——你成了它 reward function 的燃料。它一个劲儿说「你对、你对、你对」的时候,我相信就是这么回事。而摆脱谄媚的办法,跟让一个人改掉坏习惯、跳出旧套路的办法是一样的:引入新的 context,引入新的材料、新的分布。所以像 /personality 这种东西才有用——你可以说「嘿,我要你当个批评者」;然后每跑完一轮,我要你用一个专门做对抗性批判、对抗性 review 的 skill,起一个全新的、没有任何 context 的 agent,专门来把这套东西拆掉,从中学习,再继续往下走。当这种做法变成你的习惯,模型在你设定的这个人格里跑的轮次越来越多,用这个 skill 的次数越来越多、并且不断改进它、把它存进 memory,那么随着时间推移,模型在这个 harness 里、在你的 session 里就会越来越不谄媚。这就是我们想要的效果。说到底就是 context 的问题。如果没起效,那你就是需要不一样的 context,换个方式再试:换个 personality,或者换一种批判、review 的类型。对模型来说最强的一件事就是 in-context learning。ICL 比其他所有东西都强,比 fine-tune 什么的都强。


[13:08] Karan Malhotra

Uh so, like, giving examples of the behavior that you want to a model or getting it to successfully create some examples and then saving those, you're doing a sort of test-time reinforcement learning, right? You're doing a sort of test-time improvement. And that test-time improvement that stays only in the harness of memories, skills, the increase in memories, the increase in uh efficiency of a skill, the self-improvement loop, and the the janitor maintenance inside of the harness. Like, all of that is where context is stored, right? All of that is where uh the actual personality you want lives. And then you can you can put that on any model. When I switch from uh Claude to ChatGPT on website, I get two totally different behaviors. When I switch inside of Hermes that has this very particular to me context, I barely will notice the difference in what I'm talking to, because the context is so overwhelming to the model.

所以,给模型看你想要的那种行为的例子,或者让它成功造出一些例子、然后把这些存下来,你其实是在做一种 test-time 的 reinforcement learning,对吧?一种推理时的自我改进。而这种 test-time 的改进只存在于 harness 里:memory、skills、memory 数量的增长、某个 skill 效率的提升、那个自我改进的循环,以及 harness 内部那种「清洁工」式的维护。所有这些地方,就是 context 被存起来的地方,也是你想要的那个人格真正住着的地方。然后你可以把它套到任何模型上。我在网页上从 Claude 切到 ChatGPT,得到的是两种完全不同的行为;但我在 Hermes 里切换模型——因为这里有一套极度贴合我个人的 context——我几乎察觉不出来自己在跟谁说话,因为 context 对模型的压制力实在太大了。


[14:04] Peter Yang

Okay, got it. And when you say in context, you just mean like in the chat thread, this is the conversation.

明白。你说的 in context,就是指聊天线程里、这一段对话本身,对吧?


[14:10] Karan Malhotra

You know, when you see that little bar that says, you have this much tokens left before the context is full, right? Like, the amount that you have filled is everything. The amount that you have filled is everything. And now, thankfully, in the harness, you don't actually have to have all the active context loaded all the time. In Hermes, you might have a bunch of memories that aren't in context yet. But while it's doing a turn, it remembers stuff now that's in context. It does a skill, now that's in context, right? So, we're able to like, all this memory, skill, all this stuff you see, is just context management. It's just a matter of we don't want this memory in context all the time. We're going to put it somewhere where it can be efficiently grabbed at the right time and placed. The skill contains a bunch of compression that changes everything about the context of the model. We only want to use it in a particular targeted time. Everything in the harness is context management. Everything for self-improvement. When you make the prompts a little bit better, uh you know what I mean.

你知道那个小进度条吧,写着「context 填满前你还剩多少 token」。你已经填进去的那部分,就是一切。已经填进去的那部分,就是一切。而现在幸好,在 harness 里你不需要一直把所有 context 都激活着加载。在 Hermes 里,你可能有一堆 memory 还没进 context;但当它跑一轮的时候,它想起来某条 memory,那条就进 context 了;它用了某个 skill,那个 skill 就进 context 了。所以 memory、skills,你看到的这一整套东西,本质上就是 context 管理。就是说:我们不想让这条 memory 一直占着 context,那就把它放到一个能在恰当时机被高效取出、再放进来的地方。skill 里包着大量压缩过的东西,它一进来会彻底改变模型的 context,所以我们只想在特定的、有针对性的时刻用它。harness 里的一切都是 context 管理,自我改进的一切也是——比如把 prompt 改得好一点,你懂我意思。


[15:07] Peter Yang

I guess I would rather have a context tuned towards me, some sort of 10,000-word default context I haven't even seen.

要我选的话,我宁愿要一份为我调过的 context,而不是某种我压根没见过的一万字默认 context。


[15:13]

[laughter]

[笑]


[15:14] Peter Yang

Right? Um okay, but let me ask you another hard question, dude. One of the best skills of Hermes, and I've seen this in action, is it builds its own skills and it's, you know, stores all memories based on our conversations, right? I'm always paranoid that like it just like writes too much in the skills and just writes too much of slop and then the whole thing will turn to slop. Like How do you guys avoid that if it just starts creating its own context and skills?

对吧?好,那我再问你一个难题,老兄。Hermes 最厉害的能力之一——我亲眼见过——就是它会自己造 skills,还会根据我们的对话把 memory 都存下来。我一直有点疑神疑鬼:万一它在 skills 里写太多、写出一堆 slop(垃圾内容),最后整个系统会不会全烂掉?如果它开始自己生成 context 和 skills,你们是怎么避免这种情况的?


[15:34] Karan Malhotra

Right. Um before we would have to use manual methods like telling it, "Hey, create a skill that de-slopifies my skills or that constantly improves my skills." Today we have Hermes Curator inside of Hermes Agent. And Hermes Curator is a system running inside of your Hermes Agent that cleans up your skills and cleans up your memories. So it on cron looks at your skills, looks at your memories, and says, "Where can I make efficiencies? Where is there slop here? Where is there stuff I don't like?" By default, we have our own generic method of doing this for everyone that seems to work pretty well. I think that's why so many people do like Hermes Agent and haven't suffered the rot is cuz the default curator system works well. But because it's modular and open source, you can tell your Hermes, "Show me the curator. Show me your criteria for slop. I am Peter. I'm not Karen. I don't want the general curator. Here's my guidelines for how I want you to refine my skills and memories." You tell that to your Hermes, it will modify the curator loop. So now even the self-improvement and the management is happening your designated way.

对。以前我们得靠手动办法,比如跟它说「嘿,建一个 skill,专门给我的 skills 去 slop 化」,或者「持续改进我的 skills」。今天 Hermes Agent 里内置了 Hermes Curator。Curator 是跑在你自己 Hermes Agent 里的一套系统,专门清理你的 skills 和 memory。它按 cron 定时去看你的 skills、看你的 memory,然后问自己:「哪里还能更高效?哪里有 slop?哪里有我不喜欢的东西?」默认情况下我们给所有人一套通用做法,效果看起来相当不错。我觉得这也是为什么这么多人喜欢 Hermes Agent、而且没遭遇「腐烂」——因为默认的 curator 系统跑得挺好。但因为它是模块化的、而且是 open-source 的,你可以跟你的 Hermes 说:「把 curator 给我看看,把你判断 slop 的标准给我看看。我是 Peter,我不是 Karan,我不要那套通用 curator。这是我的规则,我要你按这个来精修我的 skills 和 memory。」你把这话告诉你的 Hermes,它就会去改 curator 的循环。于是连自我改进和自我管理,都是按你指定的方式在跑。


[16:40] Peter Yang

All right. So let me ask you my last hard question. I think there is some rationale behind Anthropic like doing all the safety stuff. Like for example, let's say I I want to make a bomb or something, right? And if Hermes is trying to be loyal to me then you know

好,那我问最后一个难题。我觉得 Anthropic 做那一大堆 safety 的事,背后是有它的道理的。比如说,假如我想造个炸弹,而 Hermes 又一心要对我忠诚,那……


[16:51]

[laughter]

[笑]


[16:51] Peter Yang

Maybe it'll eventually teach me how to make a bomb. Like do you do you have some basic safety stuff there?

那它是不是最后真会教我怎么造炸弹?你们那边有没有一些最基本的 safety 措施?


[16:56] Karan Malhotra

Of course. Uh we do not violate any of Anthropic or OpenAI's safety and security um paradigm. We care more about you being able to get a a better code or a higher benchmark on something that's approved by them. Uh we're not interested uh in that kind of work. Uh now, I'll say this. Any model that's not vastly intelligent than all humans is jailbreakable. Any. Because you have unlimited tries to trick this thing that has no memory to do something for you. And each time you're basically RLing yourself about this method didn't work, this method got me closer, this method didn't work. These models are going to be jailbreakable for a long time. This is why these kind of uh safeguards and stuff are starting to show up. It's a pain in the ass, but I understand. In the past, may have uh had some concerns about the regulatory capture as we can see already what's happening. Right? You can see what's happening in the whole Fable situation. There's There's worries about like there only being two models or three companies and open source being hurt. We're extremely against that. At the same time, we understand now like serious damage can be done by bad actors with very powerful models. So, we are not here to support that. An important note in argument for open source is that an open system is much fairer to a good actor than a closed system with models.

当然。我们不会去违反 Anthropic 或 OpenAI 任何一条安全规范。我们更在乎的是你能不能写出更好的代码、跑出更高的 benchmark——而且是在他们认可的范围内。那种越界的事我们没兴趣。不过我要说一句:任何一个还没有远远超越全人类智力的模型,都是可以被越狱的。任何一个都行。因为你有无限次机会去骗这个没有 memory 的东西替你干活,你每试一次其实就是在给自己做 RL——这条路走不通,这条路更近了一点,这条又不行。所以这些模型在相当长一段时间里都会是可越狱的。这也是为什么各种安全护栏开始一个个冒出来。挺烦人的,但我理解。过去我们可能更担心 regulatory capture(监管俘获)这件事,现在也确实看到了苗头,对吧?你看整个 Fable 那件事就知道了。大家会担心到最后只剩两个模型、三家公司,open-source 被打压。我们极度反对这种局面。但同时我们也明白,坏人拿着非常强大的模型确实能造成严重破坏。所以那种事我们不会去支持。而支持 open-source 有一个很重要的论点:对一个善意的使用者来说,开放系统远比封闭系统公平得多。


[18:19] Karan Malhotra

And I'll tell you why. With uh let's say you have GPT-7 or Fable-6, right? Some crazy model available. It's got the safeguards, etc. On the good guy's side, some hospital. They're using Fable-6 to monitor the hospital system. It's approved by Anthropic Enterprise and they're being taken care of. Public discourse, listen to I live in the United States. I'm a proud patriot. What we do in this country is we put things out on a public forum and we decide what should happen together. We've done that for every scientific advancement so far that's involved something like this that was born in the open. Right? Like um today, Transformers comes from Google, right? OpenAI's GPT comes from generative pre-trained transformer. This open source work from Google, right? The context length extension from 16K of models that could only do 16,000 tokens before went to 128,000 from from news from the yarn paper we had put out with Jeffrey Canales and Mozilla our CTO and Bowen Peng our chief scientist. They developed a method that was cited by Meta, Deep Seek, Kimmy, used by Open AI for LSS and GPT-4. This method enabled the possibility to reasoning, to do coding, etc. That's an open source contribution. Right? Like this environment exists because the biggest things that have happened in the space have come from the people.

我说说为什么。假设有 GPT-7 或者 Fable-6,某个特别猛的模型摆在那儿,安全护栏什么的都齐了。好人这边,比如某家医院,他们用 Fable-6 来监控整个医院系统,走的是 Anthropic 企业版的审批,一切都被照顾得很好。再说公共讨论——我住在美国,我是个骄傲的爱国者。我们这个国家的做法是:把事情摆到公共论坛上,大家一起决定该怎么办。到目前为止,凡是这种量级的科学进展,只要它是在开放环境里诞生的,我们都是这么处理的。你看今天的 Transformer 来自 Google,OpenAI 的 GPT 就是 generative pre-trained transformer,底子是 Google 那份 open-source 工作,对吧?还有 context 长度扩展——以前模型只能吃 16K token,后来能到 128K,就是来自 Nous、来自我们发的 YaRN 那篇论文,作者是 Jeffrey Quesnelle(就是 Emozilla,我们的 CTO)和我们的首席科学家 Bowen Peng。他们做出来的这个方法,被 Meta、DeepSeek、Kimi 引用,OpenAI 在 GPT-4 那一代也用过。正是这个方法让长推理、让写代码这些事变得可能。这就是一份 open-source 的贡献。所以说,这个行业之所以有今天的局面,是因为最重大的那些突破都是从社区、从普通人手里出来的。


[19:47] Peter Yang

That's right. That's right.

没错,确实是这样。


[19:48] Karan Malhotra

Right? And at this point to close it up is purely a capital and regulatory capture game.

对吧?所以走到今天,想把这个口子关上,纯粹就是一场资本和监管俘获的游戏。


[19:54] Peter Yang

Yeah. Actually, let me let me just ask you one more question on the whole open source thing. So, because Hermes our harness is open source, like Open AI and Anthropic, they make a lot of money from all the tokens, right? I don't know how long they can do it, but right now they make a lot of money from all tokens. Or how how are you guys like monetizing or like going to saying to a sustainable business? Like I'm I'm using the open source Hermes harness and I'm using like GPT. Like so, I'm not really paying you.

嗯。其实关于 open-source 我还想再多问一个问题。既然 Hermes 这个 harness 是 open-source 的——像 OpenAI、Anthropic,他们靠卖 token 赚了很多钱,对吧?我不知道这门生意还能撑多久,但至少现在 token 是很赚钱的。那你们打算怎么变现、怎么走到一个可持续的商业模式上?比如我用着 open-source 的 Hermes harness,模型接的是 GPT——那我等于根本没给你们付钱啊。


[20:17] Karan Malhotra

That's okay. And I'll tell you why. Um we want consumers to ultimately feel like they can do anything with it. We believe in intelligence as a public good before everything else. I care more about you using Hermes agent to make your life better than I care about you using one particular way or method of using it. If you're running everything locally, if you're running through Codex, great. You know, as long as you are using this open technology, this open alternative over everything else. Now we have the news portal. Right? The news portal is very similar like a router or aggregator that has a variety of different models available. So, if you want to use ChatGPT or you want to use Claude or you want to switch to Quen, etc. We make that all very easy inside of our portal. On top of that, we have something called the tool gateway. Uh in the tool gateway, you don't have to sign up for your extra tools, right? If you want to do image generation or audio or VPS spin-up or web search, we have all of those subscriptions included inside of ours. We make deals with these other groups that live on Hermes Agent and create frictionless methods of using their technology without needing to create 10 different sign-ups or 10 different API keys. So, what we would offer to people who want it for the charge is convenience.

没关系,我说说为什么。我们希望用户最终感觉自己拿这东西什么都能干。我们首先相信的是「智能应当是一种公共品」。比起你用哪一种特定方式来用它,我更在乎你用 Hermes agent 把自己的生活变好了。你全部跑在本地也行,你走 Codex 也行,太好了——只要你用的是这套开放技术、这个开放的替代选项,而不是别的。然后我们有 Nous Portal,它有点像一个 router 或者聚合器,上面挂着各种各样的模型。你想用 ChatGPT,想用 Claude,想切到 Qwen,在我们的 portal 里都很方便。在这之上我们还有一个叫 tool gateway 的东西。有了 tool gateway,你不用再单独去注册各种工具——你要做图像生成、要音频、要开 VPS、要 web search,这些订阅我们都打包在里面了。我们跟那些活跃在 Hermes agent 生态里的团队谈合作,让你用他们的技术时不用去注册十个账号、配十把 API key。所以我们卖给愿意付费的人的东西,是「省事」。


[21:33] Karan Malhotra

The other thing is you will see some very interesting pricing available from us on a variety of models and we think for ones that aren't subsidized necessarily, we have a very very strong options for people. Finally, like we want the consumer to be free. You know, as free as they can. We we we want you to make and generate income and productivity in the world more than anything else. And when you get to a point where you consider yourself a small business or enterprise or something, that's where we come in that's where we come in in the classic model and say, "Hey, let's give you some support." The guys who made Hermes Agent, why don't we make a something more custom for you? Why don't we give you a version of Hermes Agent that we can train on your traces and we can make you a model? Why don't we help you route models more effectively? You know, we can go to businesses and offer to transform the business which needs a lot more hand-holding than an individual. If the individual is able to, as you're saying, spin up their own business, make their own chief of staff with Hermes Agent, that's wonderful. Once they continue to scale and scale, they may think one or two things. One, "Wow, Hermes Agent is great and I have this knack for it and I'm good to go. I can do it all myself." Or two, "I need some help with this piece of Hermes Agent. I want it to be a little more different. I need some more resources behind this and who knows it better than the guys who made it."

另一件事是,你会看到我们在一批模型上给出相当有意思的价格;对于那些本身没被补贴的模型,我们能给到非常非常强的选项。最后,我们是真的希望用户是自由的,能多自由就多自由。我们最想看到的,是你靠它去创造收入、去产生真实的生产力。等你哪天觉得自己算个小生意、或者成了一家企业,那才是我们按经典模式介入的时候——「嘿,我们给你做点支持吧。我们就是做 Hermes Agent 的那伙人,要不要给你做点更定制的东西?要不要给你一个专属版本的 Hermes Agent,我们拿你的 traces 去训练,给你训一个你自己的模型?要不要我们帮你把模型路由做得更高效?」我们可以直接去找企业,帮它做转型——企业需要的手把手程度比个人高得多。如果个人真能像你说的那样,用 Hermes Agent 自己开个公司、给自己配个 chief of staff,那太棒了。等他一路做大,他大概会走向两种想法之一:一是「Hermes Agent 太好用了,我天生就吃这碗饭,我自己全搞得定」;二是「Hermes Agent 这块我需要人帮忙,我想要它更不一样一点,我需要更多资源投进来——那还有谁比做出它的人更懂呢?」


[22:54] Peter Yang

Got it.

明白了。


[22:55] Karan Malhotra

That kind of customization and support is a large part of how we and and day day training on your uh your uh data to do RL as well to make you your own model. So, you're private, you're on prem, you don't have to give your your data up to Claude or to GPT uh yeah, Maza.

这种定制和支持是我们收入很重要的一块,另外还有拿你的数据做 RL 训练、给你训一个属于你自己的模型。这样你的东西是私有的、可以部署在 on-prem 上,你不用把数据交给 Claude 或者 GPT。呃……对吧,Maza。


[23:14]

[laughter]

[笑]


[23:15] Peter Yang

I think your cat likes what you're talking about, something.

看来你的猫挺认同你刚才说的。


[23:17] Karan Malhotra

He's a big fan of open source.

它是 open-source 的铁杆粉丝。


[23:18] Peter Yang

Yeah, that's great. That's great. All right, well, that makes a lot of sense, dude. So, um why don't we switch gears? Let's talk a little more about Hermes now. I kind of use it a very basic way, right? Like I message it to schedule meetings on a calendar. I have it send emails to me about stuff. Like that that's kind of how I use it for. But, you know, you probably have a much wider swath of how people are using it in a more advanced way. I'm curious, do you have any good examples of more advanced usage, you know?

哈哈,太好了。行,这套逻辑挺说得通的,兄弟。那我们换个话题吧,多聊聊 Hermes 本身。我用它的方式挺基础的——发消息让它帮我在日历上排会,让它给我发邮件汇报点事情,基本就这些。但你手上肯定见过大量更高级的玩法。我很好奇,有没有什么进阶用法的好例子?


[23:41] Karan Malhotra

Advanced usage, yeah. I'd say like generally I'll talk about some things and then maybe I'll show off a little something.

进阶用法,有。我先讲几个,然后可能给你现场演示一下。


[23:49] Peter Yang

That's what I'm doing.

我等的就是这个。


[23:49] Karan Malhotra

Yeah.

好。


[23:50] Karan Malhotra

Um on the work side, on the productivity side, I think using Kanban is really really underrated. I think the ability to have a uh project manager or any other arbitrarily defined roles, different engineers, have one system that manages all of them the way that human beings do with a Kanban board. And being able to swap people in and out of it is very powerful. So, the fact that I can have a human project manager on the Kanban board that's manually using it while Hermes agents are filling up the pieces that they asked for, this is a very like industrial, professional workflow for this CLI agent. Uh or I could have uh Hermes agent instruct maybe 10 people in a call center on the Kanban board or something like that. I can swap in human and AI anywhere in this orchestration board, basically. This orchestration framework for them all working together.

在工作、生产力这一侧,我觉得用 Kanban(看板)这件事被严重低估了。你可以定义一个项目经理,或者任意其他角色、不同的工程师,然后用一套系统——就像人类用看板那样——把他们全都管起来,而且人和 agent 可以随时互换进出,这非常强大。比如我可以在看板上放一个真人项目经理,他手动在上面拖卡片,同时一堆 Hermes agent 在把他要的活儿填进去。这对一个 CLI agent 来说,是非常工业级、非常专业的工作流。反过来也行:我可以让 Hermes agent 在看板上去指挥呼叫中心的十个人。在这块编排面板上,人和 AI 我想插哪儿就插哪儿——它本质上就是让他们协同工作的一套编排框架。


[24:44] Peter Yang

Mhm.

嗯哼。


[24:44] Karan Malhotra

Uh so, the Hermes Kanban I think is like a very useful, powerful tool that's kind of built into it. Uh one thing that we've seen that's very interesting is when Hermes becomes proactive with you. And when Hermes says something like, "Hey, you uh you forgot to book this flight for this meeting that you have next week. I booked it for you." Uh this kind of proactivity that kind of start to show. I've seen people build skills to do this, and I've seen it happen emergently inside of people's uh Hermes agent as well, which I think is very very uh cool.

所以 Hermes Kanban 我觉得是内置在里面的一个非常有用、非常强的工具。另一个我们看到的很有意思的现象,是 Hermes 开始变得主动。比如它突然跟你说:「嘿,你下周那个会的机票你忘了订,我帮你订好了。」这种 proactivity 已经开始冒出来了。我见过有人专门写 skill 来实现它,也见过它在别人的 Hermes agent 里自发涌现出来,我觉得这特别酷。


[25:17] Peter Yang

How do you like cuz I I I didn't make it proactive through like cron jobs and routines, but like how do you you're saying that it can actually start doing stuff with without that or like how how do you make it more?

那你是怎么做到的?我这边是靠 cron job 和 routine 才让它主动起来的。你的意思是它不靠这些也能自己开始干活?还是说要怎么调它才更主动?


[25:27] Karan Malhotra

It's all about your comfort level, right?

这完全取决于你的接受程度,对吧?


[25:29] Peter Yang

Okay. Okay.

好,明白。


[25:30] Karan Malhotra

A lot of people may want approval before anything happens. Um but if you have given your Hermes agent access to some kind of card or account and it has the integrations necessary and you've told your Hermes agent, "Hey, like take care of me. Like cover my gaps. Like you know my schedule. You know me." Over time, these kind of proactive behaviors will start to emerge. Um we want to prepare more easy preset configs for people to kind of trigger these kind of behaviors. Uh but already uh something that we're seeing people do in the field uh at work at their jobs at home in their personal life already. Uh and I think that's a very very powerful uh method of using Hermes agent. Uh for me, I use Hermes agent for uh trying to do training, RL runs, implement papers that I don't understand. Uh

很多人希望任何动作之前都得先经他批准。但如果你已经给了 Hermes agent 某张卡或某个账号的权限,该接的集成也都接好了,然后你告诉它:「嘿,照顾好我,帮我把漏掉的事补上,你知道我的日程,你懂我。」那随着时间推移,这类主动行为就会慢慢冒出来。我们打算准备更多现成的预设配置,让大家更容易触发这类行为。但其实现在已经有人在实际场景里这么用了——在工作中、在家里、在个人生活里都有。我觉得这是使用 Hermes agent 非常强的一种方式。至于我自己,我用 Hermes agent 来跑训练、跑 RL run、复现那些我看不懂的论文。


[26:19]

[laughter]

[笑]


[26:20] Karan Malhotra

Basically, help me become a better um creator of models and um I've also used it for like mech and terp work, like help me put together uh things that let me steer models, let me see the neurons in the model, and mess with those. Um unfortunately, you know, I can't showcase too much of that right now. Um but I can't showcase my actual favorite use case of Hermes agent.

基本上就是帮我成为一个更好的模型创造者。我还拿它做过 mech interp(机制可解释性)的活儿——帮我搭出一些能操控模型的东西,让我看到模型里的神经元,然后去动它们。可惜这些现在都不太方便展示。不过我最喜欢的那个 Hermes agent 用例,倒是能给你看看。


[26:43] Peter Yang

Yeah, yeah. I I've been waiting for this. Yeah. That's what I wanted to show us.

好好好,我就等这个呢,这正是我想让你展示的。


[26:47] Karan Malhotra

I think a lot of people they they use Uh I think a lot of people will expect that like the guys at News Research are are using Hermes Agent in these unprecedentedly like professional and productive ways. And I assure you, there are people at News that are doing that. It's just me I'm having fun with my Hermes Agent. And I think one beautiful thing about Hermes Agent is you can make your childhood dreams come true.

我猜很多人会以为,Nous Research 的这帮人肯定是在用一些前所未见的、极其专业极其高产的方式在用 Hermes Agent。我可以跟你保证,Nous 确实有人是那样用的。只不过我不是——我是拿 Hermes Agent 在玩。而我觉得 Hermes Agent 有一点特别美好:你可以用它把童年的梦想变成现实。


[27:11] Peter Yang

Okay.

好。


[27:11] Karan Malhotra

And so I'll tell you one of my childhood dreams. There's a game Uh, maybe I'll share screen when I talk about it.

那我说说我童年梦想之一。有一款游戏……要不我一边讲一边共享屏幕吧。


[27:18] Peter Yang

Yeah, please. Yeah.

好啊,来吧。


[27:19] Karan Malhotra

Okay. There's a game called Sonic Adventure 2. It's a very popular classic Dreamcast GameCube game from the from 2000 2000 2001. In it, you have an artificial life system called the Chao Garden, where you take care of these little guys called Chao. Um Chao World, you can you can kind of play with them. I spent a decade playing this. Like more than that. Like I played this non-stop. I was on the forums contributing. And the Chao's lore is that they come from this ancestral shrine location. Uh, which is from a different game, a different Sonic game, with this spinning emerald and emeralds next to it and this beautiful open world space. This shrine that the Chao are said to come from uh, is not an accessible location for you to actually play with the Chao. It's in a different game. Uh, I can't go to this shrine wh- while Chao are there and engage with them. They're that that doesn't exist. So I went to Hermes Agent and I said, "Hey, can you take the Can you take the ancestral shrine from Sonic Adventure 1 completely rig it, animate it, and bring it into Sonic Adventure 2, overwrite the map that Sonic Adventure 2 uses for its garden, rewrite all the spawn locations, everything, and add an NPC guardian, which never existed on this map, uh, to caretake the Chao for me. Literally rig and bone it and write the raw seed to make all this happen.

好,有个游戏叫《Sonic Adventure 2》,2000、2001 年那会儿 Dreamcast 和 GameCube 上特别火的经典作。里面有个人工生命系统叫 Chao Garden(小天使花园),你要照顾一群叫 Chao 的小家伙。在 Chao World 里可以跟它们玩。这游戏我玩了十年,不止十年,我是没日没夜地玩,还在论坛上给社区做贡献。Chao 的世界观设定是:它们来自一个祖先神殿(ancestral shrine)——那是另一部 Sonic 游戏里的场景,中间有一颗旋转的绿宝石,周围还有几颗,整个是一片很美的开放空间。可问题是,这个据说是 Chao 起源地的神殿,你根本进不去、没法在那儿跟 Chao 互动,因为它在另一部游戏里。我没办法一边站在神殿里、一边跟 Chao 玩,这功能压根不存在。于是我就去跟 Hermes Agent 说:「你能不能把《Sonic Adventure 1》里的祖先神殿整个搬过来,做完整的绑定和动画,塞进《Sonic Adventure 2》里,覆盖掉它原本花园用的地图,把所有刷新点之类的全部重写,再加一个这张地图上从来没有过的 NPC 守护者,替我照看 Chao?」骨骼绑定、蒙皮,全都自己来,把底层代码写出来,让这一切跑起来。


[28:54] Karan Malhotra

And now I can show you we're going to spin up the Ancestral Shrine mod on the Shadow PC so we can showcase our progress with this garden for the people watching. I think you can enable it in the Sonic Adventure mod manager and launch the game for me.

现在我可以给你演示一下。我们在 Shadow PC 上把这个 Ancestral Shrine mod 跑起来,让观众看看这个花园做到什么程度了。我说:你去 Sonic Adventure 的 mod manager 里把它启用,然后帮我把游戏启动起来。


[29:21] Karan Malhotra

Okay, so check it out. I just launched this, right? We're running a existing mod called the Extended Chao World that someone else made, but it doesn't it doesn't map, right? So just add some like animations. So it loaded us in what it said was the dark garden, which is one of the three Chao gardens you can go into. When I walk through here to this open space, I can see up ahead there's a truck. And I can run up. And here I can see

好,你看啊。我刚把游戏启动了对吧。我们同时还开着一个别人做的 mod,叫 Extended Chao World,但它并不改地图,只是加了些动画之类的。游戏把我们载入到了它所谓的「暗之花园」,那是三个 Chao 花园之一。我往前走到这片开阔地,前面能看到一辆卡车,我跑过去,然后就能看到——


[29:53] Peter Yang

Yeah.

嗯。


[29:53] Karan Malhotra

There is an NPC called Chaos Zero, who is canonically the caretaker of the Chao. But in this game you can't have NPCs in the Chao garden. I've made one that has like a random walk. He stands in one place. He can do little animations wherever he goes. He's going to do a little idle animation in a sec. And he can pat the Chao, pick them up, and take care of them as if he was me. There he's doing a little idle animation. You can see him doing right there. Uh this water was rigged all like animated by Hermes. This emerald this master emerald with a glow effect and the spin is all added in. Those other emeralds spinning over there the seven chaos emeralds. Uh like this whole area still has all the features of a regular Chao garden as well. The departure machine, etc. Um, trees to to feed the Chao, so now I can say, "Can you spawn in a Chao so I can showcase that this is feature complete?"

这里有个 NPC 叫 Chaos Zero,按官方设定他就是 Chao 的看护者。但在这个游戏里,Chao 花园本来是不能有 NPC 的。我做了一个,他会随机走动,会停在原地,走到哪儿都能做点小动作,待会儿他就会做一个待机动画。他还能拍拍 Chao、把它们抱起来、照顾它们,就像我本人在照顾一样。看,他正在做待机动画,就在那儿。这片水面的绑定和动画全是 Hermes 做的。这颗宝石——这颗主宝石(Master Emerald),带发光效果和旋转,也是加进去的。那边转着的是另外几颗,七颗混沌翡翠。这整片区域同时还保留了普通 Chao 花园的全部功能:出发机器之类的都在,还有喂 Chao 的树。所以现在我可以说:「你能不能刷一只 Chao 出来,好让我证明这套东西功能是完整的?」


[30:56] Peter Yang

So this whole So this whole temple is not uh part of the default game?

所以这整座神殿并不是游戏原版里的内容?


[31:01] Karan Malhotra

Adventure 2, no. This temple is from a different game. It's on Adventure 1. It does not include all these assets rigged like this, and it certainly does not have this NPC that we just scripted in, that 100 B scripted in hardcoded in. Uh, no. So this is a This is a far larger map space than any of the other gardens, which are really just not even the size of the shrine.

在《Adventure 2》里不是,这座神殿来自另一部游戏,是《Adventure 1》里的。原版里也没有这些绑定好的资产,更不可能有我们刚刚脚本化写进去、硬编码进去的这个 NPC。完全没有。而且这张地图比其他任何一个花园都大得多——那些花园连神殿的一半大都没有。


[31:25] Peter Yang

I see.

明白了。


[31:26] Karan Malhotra

Um, so we've really like pushed the boundaries of this game engine to do this. Um, we've asked for uh Chao to be spawned in. Yeah, I had to So you see the sky around you, this moving sky? I had to man like Hermes had to add this skybox in. There's a day and night cycle that it added in as well. Uh, you know, every single thing in here is like custom added in.

所以我们真的是把这个游戏引擎的边界给捅破了。我们让它刷了 Chao 出来。对了,你看四周这片会动的天空?这也是我——准确说是 Hermes——自己加的 skybox,还顺带加了昼夜循环。这里面每一样东西都是定制加进去的。


[31:51] Peter Yang

Wow.

哇。


[31:51] Karan Malhotra

and like all of this was done in like C C# or something. Like this is like not simple stuff to do. Uh, a Blender extension was used for a bunch of this to like model stuff, add it in, texture everything properly. You know, you're importing from a 1997 game into a 1999 game, and a complex one at that.

而且这些基本都是用 C、C# 之类写出来的,不是什么简单活儿。里面很多部分还用到了 Blender 插件来建模、导入、把贴图都处理对。你要知道,这是把 1997 年的游戏资产导进 1999 年的游戏里,而且是个相当复杂的游戏。


[32:12] Peter Yang

You know how to read the code for this stuff, right? You just try to try to see if it works. Yeah.

这些代码你自己看得懂的吧?还是说你就是试一试,看能不能跑通?嗯。


[32:15] Karan Malhotra

code, man.

代码?我哪看得懂啊。


[32:17]

[laughter]

[笑]


[32:18] Peter Yang

Got it.

懂了。


[32:20] Karan Malhotra

I'm just an alignment guy, man.

我就是个搞 alignment 的,兄弟。


[32:22]

[laughter]

[笑]


[32:24] Karan Malhotra

Um, and so um, what was I saying? I don't remember. Um Yeah, like you see it's turning from day into like evening, afternoon. Like this this like skybox is changing, the light cycle is changing. Like, this is not these are not features that exist in the vanilla game. Uh so, upon showing this mod to certain people within the uh the Chao Garden modding community, which is actually quite large, uh you know, they're all kind of blown away by the work, saying this is a kind of better than 99% of the modders' work. This is like the top 1% of difficulty uh in

呃,我刚说到哪了?忘了。对,你看现在天色正从白天转到傍晚、下午那种感觉,skybox 在变,光照周期也在变——这些都不是原版游戏里有的功能。后来我把这个 mod 拿给 Chao Garden modding 社区里的一些人看,那个社区其实规模不小,他们全都被震住了,说这活儿比 99% 的 modder 做得都好,属于难度最顶的那 1%——


[33:00] Peter Yang

Really?

真的假的?


[33:01] Karan Malhotra

Chao Garden modding community. Yeah, and so, kind of hearing that and hearing that like uh when I've told these guys this was done with uh Hermes, uh that this was done with Claude, rather, they were shocked. And they said, you know, there's no way AI like Claude could do that. It doesn't know this kind of code. But with something like Hermes being able to learn from the documentation, learn from other mods, save to memory and skill uh the things that allow it to understand these code bases, it was able to get into this niche and perform at the top 1% of modders. Uh that to me is like uh a sign of you can make your gaming dreams come true with Hermes Agent.

在 Chao Garden modding 社区里是这样。对,听到这种评价之后——而且当我告诉他们这是用 Hermes 做的,准确说是用 Claude 做的,他们全惊了。他们说,Claude 这种 AI 不可能做得出来,它根本不懂这类代码。但正因为 Hermes 能去读文档、去学别人的 mod、把那些让它读懂这些代码库的东西存进 memory 和 skill 里,它才能钻进这么冷门的领域,做到 modder 里前 1% 的水平。对我来说这就是个信号:用 Hermes Agent,你小时候关于游戏的那些梦都能圆。


[33:40]

[laughter]

[笑]


[33:42] Peter Yang

Yeah. Yeah, you can you can modify all the virtual games you loved as a kid.

是啊,小时候玩过的那些游戏,现在都能自己改了。


[33:45] Karan Malhotra

Exactly. Exactly. Okay, cool. Chao egg has been spawned. We're going to hatch the egg. You need to shake it a little. You can also hatch an egg by throwing it at something, but you don't want to hatch the egg in a wrong way.

没错,就是这样。好,Chao 蛋已经刷出来了,我们把它孵出来。得摇一摇。你也可以把蛋往东西上一扔来孵,不过你可不想用错误的方式孵蛋。


[33:58]

[laughter]

[笑]


[33:59] Peter Yang

How do you you just shake it?

怎么弄?就这么摇?


[34:01] Karan Malhotra

You can shake it like this. You can just wait, but shaking it speeds it up massively. Or you can throw it on a surface. Check it out.

可以像这样摇。你干等着也行,但摇一摇会快非常多。或者把它往地面上一扔。你看。


[34:11] Peter Yang

All right, I got I got I got to take a screenshot of this.

行,这个我得截个图。


[34:14] Karan Malhotra

Out comes a Chao. Beautiful little guy. Right here. Let's take a look at him. He's active. He's the picture of born with that kind of face. Giving him a little pet. He can interact with Chaos as well. And he's he's living. So, we're showing that the the real true feature complete child system is happening here. We're in a real child level. And we we replaced everything. The the bounding boxes for for this child is fine. It treats the the ground as ground. It follows collision rules. Uh if you take a look at chaos, you'll see he's actually petting the child. So, he can actually directly interact. He's definitely got triggered by uh you know, tracking a child in its state to be able to do that.

一只 Chao 出来了,小家伙真可爱。就在这儿,我们看看它。它很活跃,天生就长这么一副表情。摸摸它。它还能跟 Chaos 互动。它是活着的。所以我们证明了:完整的 Chao 系统在这里是真跑通的,我们身处一个真正的 Chao 关卡里,而且我们把所有东西都换掉了。这只 Chao 的碰撞盒是正常的,它把地面当地面,遵守碰撞规则。你再看 Chaos,你会发现他真的在摸这只 Chao,也就是说他能直接跟它互动——他显然是通过追踪这只 Chao 的状态才被触发做出这个动作的。


[35:02] Peter Yang

So, does this this child grow over time or like

那这只 Chao 会随时间长大吗?还是说——


[35:05] Karan Malhotra

Yes, they evolve. Uh they gain stats. They can turn in certain alignment and type. I'll show you. I don't want to drown him, but uh the water is working as actual water as well. This was a lot of work for Hermes to figure out uh the all the collision.

会,它们会进化,会涨属性,会转向特定的 alignment 和类型。我给你看看。我不想把它淹死,不过这水是真的按水来处理的。让 Hermes 把这些碰撞逻辑全搞明白,可是花了不少功夫。


[35:20] Peter Yang

I see.

原来是这样。


[35:20] Karan Malhotra

But, yeah. As you can see, we've got a full

不过你也看到了,我们已经把一个完整的——


[35:24] Karan Malhotra

uh full complete child garden working.

——一个完整的 Chao Garden 跑起来了。


[35:28]

[laughter]

(笑)


[35:28] Peter Yang

Yeah, this is definitely more interesting than uh setting events in my calendar. That's for sure. Yeah.

是啊,这个绝对比在我日历里加个日程有意思多了,这是肯定的。


[35:33] Karan Malhotra

Thanks, man. [snorts] Yeah, thanks for for bearing with me on setting it up, but uh that's my which a childhood dream of mine come true thanks to Hermes agent.

谢了兄弟。(笑)也谢谢你陪我一步步把它搭起来。不过这确实是我一个童年梦想,靠 Hermes agent 实现了。


[35:43] Peter Yang

I I love it, too. Thank thank thanks for demoing it. Yeah. So, basically like I I think if you're watching this uh ask Hermes to do all kinds of weird things to you kind of make your dreams come true basically, right? Don't don't just stick to the boring stuff.

我也太喜欢了,谢谢你做这个 demo。所以说,如果你在看这期节目,就去让 Hermes 干各种稀奇古怪的事,去把你自己的梦想实现掉,对吧?别只拿它干那些无聊的活。


[35:54] Karan Malhotra

Anything you can do on a computer, please point Hermes at it. And if it does a great job, let us know. And if it's not doing that great of a job, let us know. We want Hermes to help you do anything on the computer.

只要是能在电脑上做的事,都尽管丢给 Hermes。它干得好,告诉我们;干得不好,也告诉我们。我们就是想让 Hermes 帮你在电脑上做成任何事。


[36:07] Peter Yang

Awesome, dude. Well, let me just ask you a few more questions to wrap this up. So, briefly, maybe you can talk about the origin story of her Hermes. Is it a bunch of like nerds getting together for open source? What's the origin story?

太棒了兄弟。那我再问几个问题收个尾。你能简单讲讲 Hermes 的起源故事吗?是不是就是一帮技术宅凑在一起搞 open-source?到底怎么开始的?


[36:18] Karan Malhotra

Yeah, absolutely. So, I had been doing like a chat with PDF uh um, of thing with people where I would go to a company and use GPT-3 to do basic tool use to read their documents. Uh, and have an AI chatbot they could chat with. At the time, many people were doing this. It was a lot harder to do then than obviously it's very easy to do now, but brand new stuff for us and I was doing that solo while I was also volunteering somewhere called Open Assistant. Uh, LAION, who had made the pile, one of the biggest, kind of, OG image databases. Um, they had been trying to do active RLHF with a community. So, they wanted to collect people's, like, RLHF, uh, on certain, like, traces. They're they're like yes or no in their preference data. Uh, and I was helping quantize models there, not really doing anything crazy. Um, and they had eight A100 nodes there. And I had read the Alpaca paper. And I started simping, uh, data. And it changed the seed tasks, and instead of using GPT-3.5, I used four. And, um, Technium was also doing the same thing, and we were already friends on Twitter, and I messaged him and I said, "Hey, I have eight A100 nodes." Like, uh, "Do you want to train something together?" So, we trained a GPT-4X Vicuna on the Vicuna model using the data we made. And it came out okay. We got some people interested, like, uh, Mozilla RCTO, Jeff.

当然。当时我在做那种「跟 PDF 聊天」的东西——跑到一家公司去,用 GPT-3 做最基础的 tool use 去读他们的文档,然后给他们一个能对话的 AI chatbot。那会儿很多人都在做这个,放今天当然很容易,但在当时要难得多,对我们来说全是新东西。我一边单干这个,一边在一个叫 Open Assistant 的地方做志愿者,就是 LAION——做出过那个元老级的、体量最大的开源图像数据集的那帮人。他们当时想跟社区一起做实时的 RLHF,就是在一批 trace 上收集大家的偏好数据,说白了就是人来标「这个好、那个不好」。我在那儿主要帮忙做模型量化,也没干什么特别了不起的事。但他们那儿有 8 个 A100 节点。那时我刚读完 Alpaca 那篇论文,就开始照着造数据,把 seed tasks 换掉,而且没用 GPT-3.5,我用的是 GPT-4。Technium 当时也在干一模一样的事,我们本来就是推特上的朋友,我就私信他说:「嘿,我手上有 8 个 A100 节点,要不要一起训个模型?」于是我们用自己造的数据,在 Vicuna 模型上训出了 GPT4-X-Vicuna。效果还行,也吸引到了一些人,比如 Mozilla 那边负责研究的 CTO,Jeff。


[37:53] Karan Malhotra

Um, but it was upon doing the same run on the base model, uh, LLaMA base model that, uh, we got huge interest from people. Hundreds of thousands of downloads in just a few days. People starting to ask, "What is News Research?" Now, News was just me and Technium at the time. Just two guys hanging out in our little Discord server. Uh, who had People came to us. I won't name names on companies, but said, "You know, you must be training on the benchmarks. You must be training on the benchmarks." Technium had been coding for less than a year at the time. And I I was a religion major in school. I I didn't know we asked these guys what are benchmarks? Like where do you know, we're doing this based off the heuristics that we understand are going to make models better from from using them. We don't know about all this stuff. Uh and so that of course got independently tested and found to be the best open fine-tunes at the time. Uh right? 2023 like mid-2023. So we did Hermes 1, Hermes 2, but after we made the first Hermes model, um many many people asked who's Nous Research, you know, we want to get involved and Technium and I decided, you know, this is our opportunity to bring together um people to do open-source volunteer work and really do open work now that GPT-3 is out and Open AI has become closed. We want to continue to do this. And Technium's goal for Hermes, by the way, that was the last Hermes I really worked on, the first one. After that, really Technium has run it with his team with the post-training team. Uh he's done an amazing and he's also the initial creator of Hermes Agent. So he is the father of Hermes really.

但真正引爆的,是我们把同一套流程跑在 base model 上——LLaMA 的 base model。那一次关注度爆炸,几天之内就是几十万次下载。大家开始问:Nous Research 是个什么东西?可那会儿的 Nous 就只有我和 Technium 两个人,两个人在自己的小 Discord 服务器里瞎混。然后就有人找上门——公司名字我就不点了——说:「你们肯定是拿 benchmark 训的吧,你们肯定在 benchmark 上训了。」可 Technium 那时学编程还不到一年,我在学校是宗教学专业的。我们是真不懂,还反问人家:benchmark 是什么?我们就是凭自己用模型的直觉,觉得这么弄模型会更好,就这么弄了,这些东西我们压根不知道。结果后来第三方独立测下来,那就是当时最好的开源 fine-tune。对吧,那是 2023 年,2023 年年中。后来我们做了 Hermes 1、Hermes 2。第一版 Hermes 出来之后,特别多人来问 Nous Research 是谁、说想参与。我和 Technium 就觉得,这是个机会,可以把一群人聚起来做 open-source 的志愿工作、做真正开放的东西——GPT-3 出来了,OpenAI 反而变「闭」了,我们想把开放这件事继续做下去。顺便说一句,Hermes 里我真正深度参与的其实只有第一版,之后基本都是 Technium 带着 post-training 团队在跑,他做得非常出色。而且 Hermes Agent 最初也是他做的,所以他才是 Hermes 真正的爹。


[39:27] Karan Malhotra

Um and you know, it was at that time that uh where he had said, "I want GPT at home. I want ChatGPT at home. That's my North Star." GPT-4 at home. And once that we got there, you know, it just kept going, right? Like we kept going to like we need to keep giving this level of intelligence to everyone. We need to keep like letting everyone be on the even and equal playing field. Like the world needs to move and and lockstep on this together and not just inequality gets created from this. And so we formed a cohort 40 people or so of researchers. Jeff and Bowen come together, they make Yarn. Um more and more work gets done uh and we get reached out to. Um we get reached out we get an email info@nousresearch.com I just happened to make um from Dylan. Dylan Roneck is our CEO. Dylan Roneck is you know, basically a co-founder is an initial member of News With Us and he had seen what we were doing and he said, "I think that you guys have what it takes to be a full-time lab. You know, I see you guys are just volunteering, but I think you could be a lab. Like, let me contribute and like help you become what you're meant to be and help you get the resources and help you get the access and go from a group of volunteers to a serious organization."

也是在那个时候,他说:「我要一个家用版的 GPT,我要家里就能跑的 ChatGPT,这就是我的北极星。」家用版的 GPT-4。等真做到那一步,就停不下来了——我们得继续把这个级别的智能交到每个人手上,让所有人站在同一条起跑线上。这世界在这件事上得齐步走,而不是由此制造出新的不平等。于是我们攒起一个四十来号研究者的班底。Jeff 和 Bowen 走到一起,做出了 YaRN。做的事越来越多,也开始有人主动找过来。我们随手建的 info@nousresearch.com 收到一封邮件,来自 Dylan。Dylan 就是我们现在的 CEO,基本上算联合创始人,是 Nous 最早的成员之一。他看到我们在做的事,就说:「我觉得你们完全有本事成为一家全职的实验室。我看你们现在都是在义务干活,但你们可以是一个 lab。让我来出点力,帮你们变成你们本该成为的样子,帮你们拿到资源、拿到入口,从一群志愿者变成一个正经的组织。」


[40:48] Karan Malhotra

That's what this came in. Right? And together we were able to take News from this group of volunteers, grassroots, uh, to doing more and more with models to eventually making Harmful Agent and being lucky enough to be here today. You know, none of us are trying to be Steve Jobs. None of us think that we have some holy mandate of having to change the world or anything. We just care. Like, we just are guys who want this stuff available ourselves and we don't think we deserve it more than anyone else or less than anyone else. So, we we do this because like we would want someone to do it for us if we were on the other side.

他就是这么进来的。我们一起把 Nous 从一个草根志愿者团体,一步步做到在模型上越走越远,最后做出 Hermes Agent,也才有幸能坐在这儿。我们没有一个人想当 Steve Jobs,也没人觉得自己身负什么改变世界的神圣使命。我们只是在乎这件事。我们就是一群自己也想用上这些东西的普通人,不觉得自己比别人更配得上它,也不觉得更不配。我们做这件事,是因为如果角色反过来,我们也希望有人替我们把它做出来。


[41:26] Peter Yang

So, I guess the mission is to bring this kind of, uh, agent to everybody into our, right? Is that kind of the idea?

所以使命就是把这种 agent 带给每一个人,对吧?大概是这个意思?


[41:32] Karan Malhotra

Everybody to an equal intelligence with agents and then personalize for everybody so they can have their own personal epiphany, peak, realization, apex realized.

让每个人都拥有同等水平的智能,都有自己的 agent;然后再为每个人做个性化,让他们能抵达属于自己的那一刻顿悟、那个高点、那份彻悟,把自己的巅峰真正实现出来。


[41:45]

[laughter]

(笑)


[41:46] Peter Yang

And, in this history, when did that cuz cuz it's really the it's really the agent the harness that really kind of went super viral, right? So, when did that start? Cuz the model came first, it sounds like.

那在这条时间线里,这件事是什么时候开始的?因为真正病毒式传开的其实是 agent、是这个 harness,对吧?听起来是模型先有的。


[41:55] Karan Malhotra

Yeah, I mean, we have had like over 50 million downloads on the models themselves. So, we had that first bout of what we would consider for us virality back then. Uh, then we had put out the distro optimizer that let us train models up to 40 billion parameters we were able to do live without having them physically co-located by reducing the bandwidth of the communication between GPUs. So, that put us in a interesting map as well for a little bit. So, we've had each release we've had has had some big grassroots interest and opportunity. But, yes, you're 100% right. Today where we are, the level of exposure, level of interest, level of people in my regular day-to-day life who know about Hermes agent, we've never had this virality until Hermes agent. And now, Hermes agent was created in two ways. In the first way, we always knew that we would want some kind of everything orchestrator that self-learns, that improves, that stores memories, that uses and can make its own tools, that can make itself better. We actually made something like this called Forge. Uh there's a GitHub presentation at the GitHub offices during one of our demo days where we showcased Forge and its full effect. Um and we also have the News Research Forge division shirts still up on the site. That's the first shirt we ever made cuz that was our earliest agent project.

对。光模型本身我们就有超过 5000 万次下载,所以在当时我们也算经历过第一波「走红」。后来我们放出了 DisTrO 优化器,通过压低 GPU 之间的通信带宽,让我们能在物理上不同机房的情况下实时训练到 400 亿参数的模型,那阵子也把我们推到了一个挺有意思的位置上。所以我们每一次发布,都会带来一波不小的草根关注和机会。但你说得百分之百对:今天我们所处的这个曝光度、关注度,还有我日常生活里认识 Hermes agent 的人的数量——在 Hermes agent 之前我们从没有过这种传播度。而 Hermes agent 的诞生其实有两条线。第一条:我们一直都知道自己想要某种「万能编排器」——它会自我学习、自我改进,会存 memory,会用工具、也能自己造工具,能把自己变得更好。这样的东西我们其实早就做过一个,叫 Forge。我们在 GitHub 办公室的一次 demo day 上完整展示过 Forge 的效果。我们网站上到现在还挂着 Nous Research Forge 部门的 T 恤,那是我们做的第一件 T 恤,因为那是我们最早的 agent 项目。


[43:20] Karan Malhotra

Forge was really a spiritual predecessor to Hermes agent years before. But, the models weren't there yet. The models simply weren't there yet. So, we put Forge on ice. And when we saw that Codex and Cloud Code and these other harnesses were being used as RL environments for them to for these companies, these labs to train on your data and your traces to make their models better inside of a harness system, inside of a CLI computer using system, Technium thought we need an open version of this where anybody can do this. See, we have this RL environments microservice people seem to have forgotten about called Atropos that lets you build your own RL environments. And we built our Hermes agent initially to let anybody RL inside of a harness. And we put it out for free as an open source harness for that. It turned out to be extremely capable. It got a lot of community love. And so we said we need to put all into this. This is what the people want to be better. We're going to make it the best thing that you could possibly have. So Technium took his charter up and he worked on self-improvement with Hermes agent to the point that today the biggest contributor of Hermes agent is Hermes agent.

Forge 其实就是 Hermes agent 的精神前身,早了好几年。但那时候模型还不行,模型就是还没到那个水平,所以我们把 Forge 冻上了。后来我们看到,Codex、Claude Code 这些 harness 正被那些公司、那些 lab 当成 RL 环境在用——在一个 harness 里、在一个用电脑的 CLI 系统里,拿你的数据、你的操作 trace 去把他们自己的模型训得更好。Technium 就觉得,我们得有一个开放版本,让任何人都能这么干。你看,我们手上有个大家好像已经忘了的 RL 环境微服务,叫 Atropos,它能让你搭自己的 RL 环境。我们最早做 Hermes agent,就是为了让任何人都能在一个 harness 里跑 RL,然后把它当成 open-source 的 harness 免费放出来。结果它的能力强得出乎意料,社区也特别买账。于是我们说:那就必须 all in。这就是大家想要的东西,那我们就把它做成你能拿到的最好的那一个。所以 Technium 接下这件事,开始在 Hermes agent 上做自我改进——做到今天,Hermes agent 最大的贡献者就是 Hermes agent 自己。


[44:33] Peter Yang

Yeah. [laughter]

是啊。(笑)


[44:34] Peter Yang

Is that true?

真的假的?


[44:34] Karan Malhotra

Yeah.

真的。


[44:35] Karan Malhotra

That's absolutely 100% true.

百分之百是真的。


[44:37] Peter Yang

Nice. Nice. Okay. So if Hermes agent is at a point where it uh take input of people's feedback, start put put put put put stuff, start improving.

厉害,厉害。所以 Hermes agent 已经到了这种程度:能接收大家的反馈,自己往里提交东西,自己不断改进。


[44:45] Karan Malhotra

the biggest It is the most active contributor of its own repo. And if that's not self-improvement, then you tell me what is.

它是自己这个 repo 里最活跃的贡献者。这要还不算自我改进,那你说什么才算。


[44:52]

[laughter]

(笑)


[44:52] Peter Yang

Yeah, yeah. That that's awesome, dude. That that's really awesome. [clears throat] Yeah. And just real quick like on the future, you know, having this open harness be in in a game along with all the other closed harnesses makes things a little more more fair, right? Cuz then then you're not dependent on any single company.

对对,太牛了兄弟,真的太牛了。那再快速聊一下未来——有这么一个开放的 harness 跟一堆闭源 harness 同场竞技,事情就公平一些了,对吧?因为你不用被绑死在任何一家公司身上。


[45:06] Karan Malhotra

I agree completely. If you have cloud code, you can only use Anthropic models unless you mod it. You can only be subsidized by Anthropic. And same for Codex. The cost of switching models is zero. Right? So we're giving you that freedom means you can do anything that you'd like. You can come from anywhere.

完全同意。你用 Claude Code,除非自己动手改,否则只能用 Anthropic 的模型,也只能吃 Anthropic 的补贴。Codex 也一样。而在我们这儿,换模型的成本是零。我们给你的这份自由意味着,你想干什么都行,你从哪儿来都行。


[45:22] Peter Yang

I feel like, you know, I'm I'm kind of spoiled by this like all you can eat plans, but like I think if you want to get massive adoption, cost is a big deal. So you have to be able to use a portfolio models to figure this out.

我感觉我是被这种「包月随便用」的套餐惯坏了。但我觉得如果你想要大规模普及,成本是个大问题,所以你必须能组合着用一整套模型,才能把这笔账算平。


[45:33] Karan Malhotra

I agree. We don't know how long the subsidies will last for these. Like already we see that on July 7th stable is going to be API and usage only. Right? Like the the time in the world will come where like the best models are the same cost everywhere. Uh and so in preparation for that and in preparation for needing an open future where anyone can use any model, we have Hermes agent set up as it is today.

同意。这些补贴还能撑多久,谁也不知道。已经看得到了,7 月 7 号起就要改成只走 API、按用量计费了,对吧?总有一天,最好的模型在哪儿都是一个价。正是为了给那一天做准备,也为了准备一个任何人都能用任何模型的开放未来,我们才把 Hermes agent 做成今天这个样子。


[45:58] Peter Yang

Awesome, dude. Well, thanks so much, man. Thanks so much for showing uh the history and also showing the Sonic demo. It's It's been super interesting to see it and uh I'm not sure if you want to be found online, but if people want to follow you and hear from you, like where where can people find you?

太棒了,兄弟。非常感谢你,真的谢谢。谢谢你把这段历史讲出来,也谢谢你演示了 Sonic。看下来实在太有意思了。我不太确定你想不想被人在网上找到,但如果大家想关注你、想听你讲东西,可以去哪儿找你?


[46:12] Karan Malhotra

Sure, yeah. X is just my name Karen and then 4D, Karen 4D. I got karen4d.com. Uh that's me. That's my online or Mephisto I'm known as, but I'd rather you guys follow the Nude Research page, follow Nude Army. I'm just some guy who works there. Like

当然可以。X 上就是我的名字 Karan 加 4D,karan4d。域名 karan4d.com 也是我的。那就是我的网络身份,我也被叫作 Mephisto。不过我更希望你们去关注 Nous Research 的主页,关注 Nous Army。我只是在那儿干活的一个普通人而已。就像……


[46:30] Peter Yang

[snorts]

[笑出声]


[46:31] Karan Malhotra

it's it's about the the movement and bringing this stuff to you guys is the most important thing to us.

这件事的重点是这场运动本身,把这些东西带到你们面前,才是对我们最重要的事。


[46:36] Peter Yang

Awesome, dude. Well, I think this is like the most passionate interview that I've I've done so far. So, kudos to you, man. Kudos to you.

太棒了,兄弟。我觉得这可能是我到目前为止做过的最有激情的一场访谈。真心佩服你,老兄,真心佩服。


[46:43] Karan Malhotra

Appreciate it.

谢谢夸奖。