Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI
频道: No Priors: AI, Machine Learning, Tech, & Startups
视频: https://www.youtube.com/watch?v=kwSVtQ7dziU
原文语言: en
统计: 共 93 轮
[0:00]
Code's not even the right verb anymore, right? [laughter] But I have to express my will to my agents for 16 hours a day. Manifest. [music] How can I have not just a single session of Claude code or Codex or some of these agent harnesses? How can I have more of them? How can I do that appropriately? The agent part is now taken for granted. Now the claw-like entities are taken for granted and now you can have multiple of them and now you can have instructions to them and now you can have optimization over the instructions. But there
现在"写代码"这个动词都已经不准确了,对吧?[笑声] 但我每天得花 16 个小时把我的意图传达给我的 agent。让它显化出来。[音乐] 我怎么才能不只是开一个 Claude code、Codex 或者别的什么 agent 框架的会话呢?我怎么才能多开几个?怎么才能恰当地做到这一点?现在 agent 这个部分已经是理所当然的了。现在这些 Claude 一样的东西已经是理所当然的了,你可以同时拥有好几个,你可以给它们下指令,你还可以对这些指令做优化。但是……
[0:24]
[laughter]
[笑声]
[0:24]
I mean this is why it gets to the psychosis is that this is like infinite and everything is a skill issue. Hi listeners, welcome back to No Priors. Today I'm here with Andre Karpathy and we have a wide-ranging conversation for you about code agents, the future of engineering and AI research, how more people can contribute to research, what's happening in robotics, his prediction for how agents can reach out [music] into the real world, and education in this next age. Welcome, Andre. Andre, thanks for doing this. Yeah, thank you for having me. Uh so it's been a very exciting couple of months in AI. Uh yeah, you could say that.
我是说,这就是为什么会让人陷入那种近乎疯魔的状态——因为这东西是无限的,而且一切问题都归结为"是你自己技术不到位"。各位听众好,欢迎回到 No Priors。今天我和 Andrej Karpathy 坐在一起,要为大家带来一场范围很广的对话,主题包括代码 agent、工程和 AI 研究的未来、怎样让更多人能参与到研究中来、机器人领域正在发生什么、他对 agent 如何延伸到现实世界的预测,以及在下一个时代教育会是什么样子。欢迎你,Andrej。Andrej,谢谢你愿意来。好的,谢谢你们邀请我。嗯,过去这几个月 AI 领域真是非常激动人心。是啊,可以这么说。
[1:03]
I remember um walking into the office at some point and you were like really locked in and I was asking what you were up to and you're like, I just I have to code for 16 hours a day or code's not even the right verb anymore, right? But I have to um express my will to my agents for 16 hours a day. Manifest um because like there's been a jump in capability. Uh what's happening? Tell me about your experience. Yeah, I kind of feel like I was just in this perpetual I still am often in this state of AI psychosis just like all the time um because there was a huge unlock in what you can achieve as a person as an individual, right? Because you were bottlenecked by, you know, your typing speed and so on. But now with these agents it really, I would say in December is when it really just something flipped where I kind of went from 80/20 of like, you know, uh to like 20/80 of writing code by myself versus just delegating to agents. And I don't even think it's 20/80 by now. I think it's a lot more than that. I don't think I've typed like a line of code probably since December basically.
我记得有一次走进办公室,看到你整个人特别投入、特别"锁定"那种状态,我就问你在忙什么,你说,我就是得每天写 16 个小时的代码——其实"写代码"这个词都不太对了,对吧?但我每天得花 16 个小时把我的意图传达给我的 agent。让它显化出来。因为能力上确实出现了一次飞跃。这到底是怎么回事?跟我说说你的体验。是这样,我感觉自己一直处在——其实现在也常常处在——那种"AI 疯魔"的状态里,简直是无时无刻不在这个状态。因为作为一个个体,你能做到的事情发生了一次巨大的解锁,对吧?因为以前你是被自己的打字速度之类的东西卡住的。但现在有了这些 agent,真的——我会说大概是在 12 月,某个东西彻底翻转了,我从原来 80/20、也就是 80% 自己写代码、20% 交给 agent,变成了 20/80。而且我觉得现在都不止 20/80 了,比那个比例还要悬殊得多。我觉得我从 12 月开始基本上一行代码都没亲手敲过。
[1:59]
[laughter]
[笑声]
[2:00]
Um which is like an extremely large uh change. Um I was talking to it like for example, I was talking about it to for example my parents and so on and I don't think like a normal person actually realizes that this happened or how dramatic it was. Like literally like if you just find a random software engineer or something like that at their at their desk and what they're doing, like their default workflow of, you know, building software is completely different as of basically December. Uh so I'm just like in this state of psychosis of trying to figure out like what's possible, uh trying to push it to the limit. How is it how can I have not just a single session of, you know, um Claude code or Codex or some of these agent harnesses? How can I have more of them? How can I do that uh appropriately? And then how can I use these claws? What are these claws? Uh and uh so there's like a lot of new things. I want to be at the forefront of it, you know, and I'm very antsy that I'm not at the forefront of it and I see lots of people on Twitter doing all kinds of things and they all sound like really good ideas and I need to be at the forefront or I feel extremely nervous. And so I guess I'm just in this psychosis of like what's possible like because it's unexplored fundamentally. Well, if you're nervous, the rest of us are are nervous. We have a we have a team that we work with at Conviction that their setup is everybody is like, you know, none of the engineers write code by hand and they they're all microphoned and they just like whisper to their agents all the time. It's the strangest work setting ever. Uh and I thought they were crazy and now I like I fully accept I was like, oh this was the way. Like you're just ahead of it. Um what uh how do you think about your own capacity now to like explore or to do projects? Like what is it limited by? Yeah, what is it limited by? Uh just I think everything like so many things even if they don't work, I think to a large extent you feel like it's a skill issue. It's not that the capability is not there. It's that you just haven't found a way to string it together of what's available. Like I just don't I didn't give good enough instructions in the agents from the file or whatever it may be. I don't have a nice enough memory tool that I put in there or something like that. So it all kind of feels like skill issue when it doesn't work to some extent. You want to see how you can parallelize them etc. and you want to be Peter Steinberg basically. Uh so Peter is famous. He has a funny photo where he's in front of a monitor with lots of uh like he uses Codex. So lots of Codex agents tiling the the monitor and they all take about 20 minutes if you prompt them correctly and use the high effort. And so they all take about 20 minutes. They have multiple, you know, 10 repos checked out. And so he's just um going between them and giving them work. It's just like you can you can you can move in much larger macro actions. It's not just like here's a line of code, here's a new function. It's like here's a new functionality and delegate it to agent one. Here's a new functionality that's not going to interfere with the other one. Give it agent two. And then try to uh review their work as best as you can
嗯,这是一个极其巨大的变化。我比如说跟我父母聊过这件事,我觉得一个普通人其实根本意识不到这件事已经发生了,也意识不到它有多么剧烈。比如说,你随便找一个软件工程师,看他坐在桌前在做什么——他构建软件的那套默认工作流,基本上从 12 月开始就完全不一样了。所以我就是处在这种疯魔状态里,努力搞清楚什么是可能的,努力把它推到极限。我怎么才能不只是开一个 Claude code、Codex 或者别的什么 agent 框架的会话?我怎么才能多开几个?怎么才能恰当地做到?然后我怎么去用这些 Claude?这些 Claude 又到底是什么?所以有一大堆新东西。我想站在最前沿,而且我一旦觉得自己不在最前沿就会非常焦躁。我看到推特上一大堆人在做各种各样的事情,听起来都像是很好的点子,我必须在最前沿,否则我会极其紧张。所以我猜我就是处在这种疯魔里,琢磨什么是可能的——因为这本质上是一片未被探索的领域。嗯,如果连你都紧张,那我们这些人就更紧张了。我们在 Conviction 有一个合作的团队,他们的配置是这样的:所有工程师都不再用手敲代码,每个人都戴着麦克风,整天就对着自己的 agent 小声说话。这是我见过最奇怪的工作场景。我本来觉得他们疯了,可现在我完全接受了,我心想,哦,原来这才是正确的方式。你只是走在了前面。那你怎么看待你自己现在的产能——去探索、去做项目的能力?它现在是被什么限制住的?是啊,被什么限制住的?我觉得几乎所有事情——很多事情即使做不成,我觉得在很大程度上你会觉得那是"技术不到位"的问题。不是说能力不在那儿,而是你还没找到把这些能力串起来的办法。比如我就是没在 agents 文件里给出足够好的指令,或者别的什么。我没有放进一个足够好用的记忆工具之类的。所以当事情做不成的时候,某种程度上感觉都像是"技术不到位"。你会想看看怎么把它们并行起来等等,你想成为 Peter Steinberg 那种人。Peter 很有名,他有一张很搞笑的照片,他坐在一个显示器前面,上面排满了——他用的是 Codex——所以是一堆 Codex agent 平铺在整个显示器上,只要你提示得当、用上高强度模式,它们每个大概要跑 20 分钟。所以它们都各跑大概 20 分钟。他同时 checkout 了好多个仓库,可能十个之类的。所以他就是在它们之间来回切换、给它们派活。这就好像——你可以以大得多的"宏操作"来推进工作。不再是"这是一行代码、这是一个新函数",而是"这是一个新功能,交给 agent 一号"。"这是另一个不会和前面那个冲突的新功能,交给 agent 二号。"然后你尽可能去 review 它们的产出。
[4:37]
[laughter]
[笑声]
[4:37]
depending on how much you care about that code. Like where are these macro actions that I can like manipulate my software repository by? And like another agent is doing some like research, another agent is writing code, another one is coming up with a plan for some new implementation. And so everything is just like happens in these like macro actions over your repository. Um and you're just trying to become like really good at it and develop like a muscle memory for it is extremely um Yeah, it's very rewarding number one because it actually works. Uh but it's also kind of like the new thing to learn. So that's why hence the psychosis. Yeah, I I do feel like my instinct is like whenever I'm waiting for an agent to complete something, the obvious thing to do is like, well, I can do more work, right? Like if I have access to more tokens then like I should just parallelize at tasks. And so that's that's very stressful because if you don't feel very bounded by your ability to spend on tokens, then you know, you are the bottleneck in the system that is max capability. Yeah, if you're not maximizing your subscription at least. And ideally for multiple agents. Like if you run out of the quota on Codex, you should switch to Claude or whatnot. I don't know. Like that's what I've been trying to do a little bit and I feel nervous when I have subscription left over. That just means I haven't maximized my token throughput. So I actually kind of experienced this when I was a PhD student. You would feel nervous when your GPUs are not running. Like you have GPU capability and you're not maximizing your the available flops to you. But now it's not about flops, it's about tokens. So what is your token throughput and what token throughput do you command? I would actually argue that it's very interesting that we had, you know, at least 10 years where in many engineering tasks people just did they didn't feel compute bound. Right? Um and now the entire industry feels that now. They feel like they they they felt resource bound uh and now that you have this big capability jump, you're like, oh, actually it's not, you know, my ability to access the computer anymore. Like I'm I'm the binding constraint. Yeah, it's a skill issue. Which is very empowering cuz um yeah, cuz you could be getting better. So that's why that's why I think it's very addictive because there's unlocks when you when you get better. Where do you think it goes? Like if you just think about like, okay, you know, Andre's iterating and everybody else is for 16 hours a day getting better at using coding agents. Like what does it look like in a year? Of like you've reached mastery.
review 多仔细取决于你对那段代码有多在意。就像,我能用哪些"宏操作"来操控我的软件仓库?还有,另一个 agent 在做某种调研,另一个 agent 在写代码,又有一个 agent 在为某个新的实现想方案。所以一切都是以这些针对你仓库的"宏操作"在发生。你就是努力让自己变得特别擅长这个,培养出一种肌肉记忆,这非常……是啊,这非常有成就感,第一是因为它真的管用。但它同时也是那种新需要学的东西。所以这也是为什么会疯魔。是啊,我确实感觉,我的本能是:每当我在等一个 agent 完成某件事的时候,显而易见该做的就是——那我可以多干点活,对吧?如果我能用更多的 token,那我就应该把任务并行起来。所以这挺让人有压力的,因为如果你不觉得自己花 token 的能力被卡住,那你就成了这个系统里的瓶颈——这个系统本来能力是拉满的。是啊,至少你得把你的订阅额度用满。理想情况下还得是多个 agent 一起用。比如你 Codex 的配额用光了,你就该切到 Claude 之类的。我也说不好。这就是我一直在尝试做的,而且当我还有订阅额度剩下时我会紧张。那只意味着我没把我的 token 吞吐量拉满。其实我读博的时候就经历过类似的感觉。当你的 GPU 没在跑的时候你会紧张。你有 GPU 算力却没把可用的 flops 用满。但现在不是关于 flops 了,是关于 token。所以你的 token 吞吐量是多少?你能调动多大的 token 吞吐量?我其实想说,很有意思的一点是,过去至少有十年,在很多工程任务里人们并不觉得自己被算力卡住,对吧?而现在整个行业都感受到了这一点。他们感觉自己被资源卡住了,而现在你有了这么一次巨大的能力跃升,你会发现,哦,其实瓶颈不再是我访问算力的能力了。是我自己成了那个约束条件。是啊,是"技术不到位"的问题。这其实很赋能,因为——是啊,因为你是可以变得更厉害的。所以这也是为什么我觉得它很让人上瘾——因为你一旦变厉害就会有新的解锁。你觉得这会走向何方?比如你想象一下,Andrej 在不断迭代,其他每个人也都每天花 16 小时让自己更擅长用编程 agent。一年后会是什么样子?也就是你已经达到"精通"的时候。
[6:48]
[laughter]
[笑声]
[6:49]
Yeah, what does mastery look like, right? At the end of the year or like two, three years, five years, 10 years, etc. Well, I think everyone is basically interested in like going up the stack. So I would say it's yeah, it's not about a single session with your agent. Multiple agents, how do they collaborate and teams and so on. So everyone's trying to figure out what that looks like. And then I would say Claude is also kind of an interesting direction because it really, when I say a Claude, I mean this like layer that kind of takes persistence to a whole new level. Like it's something that like keeps looping. It's it's like um it's not something that you are interactively in the middle of. It kind of like has its own little sandbox, its own little you know, it kind of like does stuff on your behalf even if you're not looking kind of thing. Um and then also has like maybe more sophisticated memory systems etc. that are not yet implemented in agents. So um Open Claude has a lot more sophisticated memory I would say than what you would get by default uh which is just a memory compaction when your context runs out, right? You think that's the piece that resonated for more users versus like perhaps like broader tool access? For Open Claude? Yeah. Uh there's like I think there's at least five things that are really good ideas in here. Yeah, good job, Peter. I mean Peter has done a really amazing job. Um I saw him recently. Uh and I talked to him about it and I he's very humble about it. But I think he innovated simultaneously in like five different ways and put it all together. Um so for example like the soul and D document. Like he actually really crafted a personality that is kind of compelling and interesting. And I feel like a lot of the current agents they don't get this correctly. I actually think a Claude has a pretty good personality. It feels like a teammate uh and it's excited with you etc. I would say um for example Codex is a lot more dry um which is kind of interesting because [laughter] in it's true. You know, it doesn't it and the other thing I would say is for example with Claude I think they dialed the sycophancy fairly well where when Claude gives me praise, I do feel like I slightly deserve it because sometimes I kind of give it like not very well formed thoughts and uh I give it an idea that I don't think it's fully baked and it doesn't actually react very strongly. It's like, oh yeah, we can implement that. But when it's a really good idea by my own account, it does uh seem to reward it a bit more. And so I kind of feel like I'm trying to like earn its praise which is really weird. And so I do think the personality matters a lot uh and I think a lot of the other uh tools maybe don't appreciate it as much. And I think in this aspect also Peter really cares about this and so that was correct. And then the memory system and then uh just, you know, he's just having fun with this um and then the the single WhatsApp portal to all of the automation.
是啊,精通会是什么样子,对吧?一年后,或者两三年、五年、十年后等等。嗯,我觉得每个人基本上都想往技术栈的上层走。所以我会说——重点不在于跟你的 agent 进行单次会话。而是多个 agent,它们如何协作、如何组成团队等等。所以每个人都在努力搞清楚那会是什么样。然后我会说 Claude 也是一个挺有意思的方向,因为它真的——我说的 Claude 指的是这样一个层:它把"持续性"提升到了一个全新的水平。它是那种会一直循环下去的东西。它不是你需要交互式地全程介入的东西。它有点像有自己的小沙盒、自己的小空间,它会替你做事,哪怕你没在看也照样做。然后它可能还有更复杂的记忆系统等等,这些目前还没在 agent 里实现。所以 Open Claude 的记忆系统我会说要比你默认得到的复杂得多——默认的无非就是上下文用完时做一次记忆压缩,对吧?你觉得相比于更广泛的工具访问权限,记忆这一块才是更打动用户的部分吗?说的是 Open Claude?是的。我觉得里面至少有五个真的很棒的点子。是啊,干得漂亮,Peter。我是说 Peter 真的做得相当出色。我最近见过他。我跟他聊了这个,他对此非常谦逊。但我觉得他在大概五个不同的方向上同时做出了创新,并把它们全部整合到了一起。比如说那个 soul 文档之类的。他真的精心打造了一个相当有吸引力、相当有趣的人格。我感觉现在很多 agent 没把这一点做对。我其实觉得 Claude 的人格相当不错。它感觉像是个队友,会跟你一起兴奋等等。我会说,比方说 Codex 就要干巴巴得多,这其实挺有意思的,因为 [笑声] 这是真的。它就是不会……另一件我想说的是,比方说 Claude,我觉得他们把那种"奉承"的程度调得相当到位——当 Claude 夸我的时候,我确实会觉得自己稍微配得上这份夸奖。因为有时候我给它的想法其实没成形,我给它一个我自己都觉得没想透的点子,它并不会反应特别强烈。它会说,哦对,这个我们可以实现。但当我自己觉得这真的是个好点子的时候,它确实会显得更愿意去奖励它一点。所以我有点感觉自己是在努力"挣得"它的夸奖,这真的很怪。所以我确实觉得人格这件事非常重要,而我觉得很多其他工具可能没那么重视它。在这一点上,Peter 也是真的很在意,所以这点他做对了。然后是记忆系统,然后就是——他就是在拿这个玩得很开心,再然后是那个连通所有自动化的单一 WhatsApp 入口。
[9:17]
Yeah. Is there something that you have done personally with your claws beyond software engineering that you think is fun or interesting? Yeah, so in January I had a claw I went through a period of claw psychosis. So I built um I have a claw basically that takes care of my home and I call him Dobby the elf uh claw. Um and uh basically I used uh the agents to find all of the smart home subsystems of my home on the local area network which I was kind of surprised that it worked out of the box. Like I just told it that I think I have Sonos at home. Like can you try to find it? And it goes and it did like IP scan of all of the um basically um computers on the local area network and and found the Sonos thing uh the Sonos uh, system and it turned out that there's no password protection or anything like that. It just logged in and it's like, "Oh, yeah, you have these Sonos systems installed. I Let me try to reverse engineer how it's working." It does some web searches and it finds like, "Okay, these are the API endpoints." And then it's like, "Do you want to try it?" And I'm like, "Whoa, like you just did that." And I'm like, "Yeah, can you try to play something in the study?" And, uh, it does and music comes out and I'm like, "I can't believe I just That's crazy. That's like three prompts. Yeah.
是啊。除了软件工程之外,你有没有用你的 Claude 亲自做过什么你觉得好玩或者有意思的事情?有啊,一月份的时候我有一个 Claude——我经历过一段"Claude 疯魔"期。所以我搭了一个——我基本上有一个 Claude 专门照看我的家,我管它叫 Dobby 小精灵 Claude。我基本上用这些 agent 去找出我家里局域网上所有的智能家居子系统,它居然开箱即用就成功了,这让我挺意外的。我就是告诉它,我觉得我家里有 Sonos 设备。你能不能试着找一下?然后它就去对局域网上所有的——基本上所有电脑——做了一次 IP 扫描,找到了那个 Sonos 设备、Sonos 系统,结果发现根本没有密码保护之类的东西。它直接就登录进去了,然后说,哦对,你装了这些 Sonos 系统。让我试着逆向一下它是怎么工作的。它做了一些网络搜索,找到了,好的,这些是 API 端点。然后它说,你想试试看吗?我心想,哇,你刚才就这么把它干成了。然后我说,好啊,你能不能试着在书房放点音乐?它真的放了,音乐就出来了,我简直不敢相信我刚刚——这太疯狂了。这才三句提示而已。是啊。
[10:19]
I can't believe I just typed in like, "Can you find my Sonos?" and then suddenly it's playing music. And it did the same for lights. And so like it kind of hacked in, figured out the whole thing, uh, created APIs, created dashboard so I could see the command, uh, kind of center of like all of my lights in the home. And then it was like switching lights on and off and, you know, so I can ask it like, "Dobby, it's sleepy time." And when it's sleepy time that just means all the lights go off, etc. and like so on. So it controls all of my lights, my HVAC, my shades, uh, the pool and, uh, the spa and also my security system. So I have a camera pointed outside of the house and anytime someone rolls in I have a Quinn, uh, a Quinn, uh, model that looks at the videos. So first of all there's change detection. Right.
我简直不敢相信,我刚才就输入了一句"你能找到我的 Sonos 吗",然后突然之间它就在放音乐了。它对灯也做了一样的事。所以它基本上就是黑进去、搞清楚了整套系统、做出了 API、做出了一个 dashboard,让我能看到那个指挥中心、控制我家里所有的灯。然后它就能开关灯了,所以我可以跟它说,比如,Dobby,到睡觉时间了。一到睡觉时间就意味着所有灯都关掉,等等之类的。所以它控制着我家里所有的灯、我的暖通空调、我的窗帘、泳池、温泉浴池,还有我的安防系统。我有一个朝着房子外面的摄像头,任何时候有人靠近,我有一个 Quinn 模型在看这些视频。所以首先是变化检测。对。
[10:58]
And then based on change detection it goes to Quinn and then it actually like tells me, um, it sends me a text to my WhatsApp. It shows an image from the outside and it says, "Hey, a FedEx truck just pulled up. FedEx truck just pulled up and you might want to check it and you got new mail or something like that." And Dobby just text me this. This is really incredible. Um, so so Dobby is in charge of the house. I text through with it through WhatsApp, um, and it's been like really fun to have these macro actions that maintain my house. I haven't like really pushed it, uh, like way more beyond that and I think people are doing a lot more crazy things with it, uh, but for me even just the home automation setup I used to use like six apps, uh, completely different apps and I don't have to use these apps anymore. Like Dobby controls everything in natural language. It's amazing. Um, and so I think like I haven't even pushed the paradigm fully but already that is so helpful and so inspiring I would say. Do you think that's indicative of like what people want from a user experience perspective with software, right? Because I I don't think, you know, it's pretty ignored that it takes humans effort to like learn new software, like new UI. Yeah. I think, uh, to some extent that's right. It's like working backwards from how people think an AI should be because what people have in their mind of like what an AI is is not actually what an LLM is by by like in the raw sense. Like LLM is a token generator, you know, like more tokens come out. But what they think of is like this this persona identity that they can tell stuff and it remembers it, you know? And, uh, it's just kind of an entity behind the WhatsApp. It's like a lot more understandable. Mhm. Uh, so I think to some extent it's like matching the expectations that humans already have for what an AI should behave but under the hood it's like a lot of technical details go into that. And LLMs are too raw of a primitive, uh, to actually, um, type check as AI I think for most people if that makes sense. Yeah. Um, I think that's like how we understand what the AI is and like the, um, description of it as Dobby or some persona obviously resonates with people. Um, I also think that it it uh, the unification that you did across your six different software systems for your home automation speaks to a different question of like do people really want all of the software that we have today? Yeah. Right? Um, because I I would argue like, well, you have the hardware but you've now thrown away the software or the UX layer of it. Um, do you think that's what people want? Yeah, I think there's this like there's this sense that these apps that are on the app store for using these smart home devices, etc. Uh, these shouldn't even exist kind of in a certain sense. Like shouldn't it just be APIs and shouldn't agents be just using it directly? And, um, wouldn't it like I can do all kinds of home automation stuff that, uh, in any individual app will not be able to do, right? Um, and an LLM can actually drive the tools and call all the right tools and do uh, do pretty complicated things. Um, and so in a certain sense it does point to this like maybe there's like an overproduction of lots of custom bespoke apps that shouldn't exist because agents kind of like crumble them up and everything should be a lot more just like exposed API endpoints and agents are the glue of the intelligence that actually like tool calls all the all the parts. Um, another example is like my treadmill. Uh, there's an app for my treadmill and I wanted to like keep track of how often I do my cardio, uh, but like I don't want to like log into web UI and go through a flow and etc. Like all this should just be like make APIs available and this is kind of, you know, going towards the agentic, um, sort of web or like agent first, uh, tools and all this kind of stuff. So I think the industry just has to reconfigure in so many ways that's like the customer is not the human anymore. It's like agents who are acting on behalf of humans and this refactoring will be will probably be substantial in a certain sense. One way that people sometimes push back on this is like, do people Do you Do we expect people to write code some of these tools? Do we expect normal people to do this kind of stuff that I described? Mhm. But I think to some extent this is just, you know, technology as it exists today and right now there is some write coding and I'm actually watching it and I'm working with the system but I kind of feel like this kind of stuff that I just talked about this should be free like in a year or two or three. There's no write coding involved. This is trivial. This is table stakes. This is like any AI, even the open source models, etc. can like do this. You should be able to translate it from a less technical humans intent very easily to this outcome.
然后基于变化检测,它会去调用 Quinn,接着它真的会告诉我——它会给我的 WhatsApp 发一条消息。它会附上一张外面的图片,然后说,嘿,刚有一辆 FedEx 卡车停过来了。FedEx 卡车刚停过来,你也许想去看看,你有新的邮件之类的。然后 Dobby 就把这个发给我了。这真的太不可思议了。所以 Dobby 负责管这个家。我通过 WhatsApp 跟它来回发消息,能有这些维护我家的"宏操作"真的特别好玩。我还没有真的把它推得更远,我觉得很多人正在用它做疯狂得多的事情,但对我来说,光是这套家庭自动化的配置——我以前要用六个 app,完全不同的六个 app,现在我再也不用这些 app 了。Dobby 用自然语言就能控制一切。太棒了。所以我觉得我甚至都还没把这个范式完全推到位,但光是这样就已经这么有帮助、这么有启发性了。你觉得这能说明,从用户体验的角度看,人们想要的软件就是这样的吗?因为我觉得有一点被大大忽视了,那就是人类要付出努力去学新软件、学新 UI。是啊。我觉得某种程度上是这样的。这就像是从"人们觉得一个 AI 应该是什么样"反推回来——因为人们脑子里想的那个"AI 是什么",其实并不是 LLM 在原始意义上的样子。LLM 是个 token 生成器,你懂的,就是更多 token 出来。但人们想的是这么一个——这么一个有人格、有身份的东西,他们可以跟它说事情,它会记住,你懂吧?它就是 WhatsApp 背后的一个实体。这就好理解多了。嗯。所以某种程度上,这就是在匹配人类本来就对"一个 AI 应该如何表现"所抱有的期待,但底层其实有大量技术细节。而 LLM 作为一个原语来说太"原始"了,对大多数人来说它没法真的被"类型检查"成 AI,如果你懂我意思的话。是啊。我觉得这就是我们理解"AI 是什么"的方式,而把它描述成 Dobby 或者某个人格,显然会引起人们的共鸣。我还觉得,你在自己的家庭自动化里把六个不同软件系统统一起来这件事,触及了另一个问题:人们真的想要我们今天拥有的所有这些软件吗?是啊。对吧?因为我会说,你保留了硬件,但你现在已经把它的软件、或者说 UX 层给扔掉了。你觉得那是人们想要的吗?是啊,我觉得有这么一种感觉——应用商店里那些用来操控这些智能家居设备之类的 app,从某种意义上说它们甚至根本就不该存在。难道不应该只有 API,然后让 agent 直接去用它们吗?而且——我能做各种各样的家庭自动化的事情,是任何单一 app 都做不到的,对吧?而 LLM 真的可以驱动这些工具、调用所有正确的工具,做相当复杂的事情。所以从某种意义上说,这确实指向了这一点:也许存在着对大量定制化、专门化 app 的"过度生产",这些 app 不该存在,因为 agent 基本上把它们都揉碎了,一切都应该更像是裸露出来的 API 端点,而 agent 就是那个智能的胶水,真正去工具调用所有的部件。另一个例子是我的跑步机。我的跑步机有个 app,我想记录一下我多久做一次有氧,但我不想登录某个网页 UI、走一整套流程之类的。这一切都应该只是把 API 提供出来,这就是在朝着 agent 化的网络、或者说"agent 优先"的工具这类东西走。所以我觉得整个行业必须在非常多的方面重新洗牌——因为客户不再是人类了,而是代表人类行动的 agent,这种重构在某种意义上可能会是相当庞大的。人们有时候反驳这个观点的一种方式是:人们……你们指望人们去给这些工具写代码吗?我们指望普通人去做我刚才描述的那种事情吗?嗯。但我觉得某种程度上,这只是今天这个时间点上技术的样子,现在确实有一些写代码的成分,我确实在盯着它看、在跟系统一起工作,但我有种感觉,我刚才说的这类事情——这在一两年或者两三年内应该会变成免费的、不费力的。完全不涉及写代码。这会变得很平凡。这会成为基本盘。任何 AI,哪怕是开源模型之类的,都能做到这个。你应该能很容易地把一个技术水平不那么高的人的意图翻译成这个结果。
[15:00]
Yeah. Today it's write coding and it's involved and not many people are going to do it but
是啊。今天它还是要写代码,还是挺费劲的,不会有很多人去做,但是……
[15:02]
And you still have to make some design decisions, right? We were talking about like we take frames for example. Yeah. Yeah. But I kind of feel like this will just, uh, start to the barrier will just come down and it's just ephemeral software on your behalf and some kind of like claw is handling all the details for you but you're not involved. Claw has a Claw has a machine and it will figure it out and it's just presenting you UIs and you're like saying stuff, you know? Mhm. Why haven't you, um, I guess like pushed the boundaries of what you can do personally with claws? Like is it, you know, you're focusing on more important projects, auto research, etc. or, uh, you're climbing the hill to mastery or something else, right? Yeah, I just feel like I'm so distracted by everything so I spend I [laughter] spend like a week on the claw stuff and I I have more to do almost, um, but I will say that, um,
而且你还得做一些设计决策,对吧?我们刚才在聊比如"取帧"这种事。对。是啊。但我有种感觉,这个门槛就是会——会一路降下来,最后它就成了替你临时生成的、用完即弃的软件,某种 Claude 帮你打理所有细节,但你不用参与。Claude 有一台机器,它会把事情搞定,它只是把 UI 呈现给你,你就在那儿说说话,你懂吧?嗯。你为什么没有——我猜,亲自去拓展你用 Claude 能做的事情的边界?是因为你在专注于更重要的项目,比如 auto research 之类的,还是你在攀登通往"精通"的那座山,还是别的原因?是啊,我就是觉得自己被所有事情搞得太分心了,所以我会花——我 [笑声] 花了大概一周时间在 Claude 这些事情上,我几乎还有更多想做的,不过我得说……
[15:50]
It's like Jensen told us we're all just busier, unfortunately.
就像 Jensen 跟我们说的,很不幸我们所有人都只是变得更忙了。
[15:53]
Uh, I didn't really take advantage of a lot of like email and calendar and all this other stuff and I didn't really have access cuz I'm still a little bit like suspicious and it's still very new and rough around the edges. So I didn't want to give it like full access to my digital life yet and part of it is just the security, privacy and uh, just being very cautious in that in that realm. And, um, so some of it is like held back by that I would say. Yeah, maybe that's like the dominant dominant feature but some of it is also just I feel so distracted because I feel like I had a week of claw and then other stuff is happening and What was the, um, I mean you've talked about like being able to train or at least optimize a uh, a a model as a task you want to see agents do for a long time. Like what was the motivation behind auto research? Auto research, yeah. So I think like I had a tweet earlier where I kind of like said something along the lines of to get the most out of the tools that have become available now you have to remove yourself as the as the bottleneck. You can't be there to prompt the next thing. You're You need to take yourself outside. Um, you have to arrange things such that they're completely autonomous. And the more you you know, how can you maximize your token throughput and not be in the loop? This is the this is the goal. And so I kind of mentioned that the the name of the game now is to increase your leverage. Uh, I put in just very few tokens just once in a while and a huge amount of stuff happens on my behalf. And so auto research like I tweeted that and I think people liked it and whatnot but it they haven't like maybe worked through like the implications of that and for me auto research is an example of like an implication of that. Where it's like I don't want to be like the researcher in loop like looking at results, etc. Like I'm I'm holding the system back. So the question is how do I refactor all the abstractions so that I'm not I have to arrange it once and hit go. The name of the game is how can you get more agents running for longer periods of time without your involvement doing stuff on your behalf? And auto research is just, yeah, here's an objective, here's a metric, here's your boundaries of what you can and cannot do. And go. And, uh, yeah, it worked.
我其实没怎么去利用邮件、日历那一大堆别的东西,我也没真的给它访问权限,因为我还是有点……有点疑虑,而且它还很新、边边角角还很粗糙。所以我还不想给它对我整个数字生活的完全访问权限,这部分原因就是安全、隐私,以及在那个领域非常谨慎。所以有一部分是被这个拖住了,我会说。是啊,也许这才是主要的、占主导的因素,但还有一部分原因就是我觉得自己太分心了——因为我感觉我花了一周在 Claude 上,然后别的事情又冒出来了。那个——我是说你以前谈过,能训练、或者至少能优化一个模型,是你长期以来想看到 agent 去做的一项任务。那 auto research 背后的动机是什么?auto research,是的。我觉得我之前发过一条推,大意是说,要想把现在已经可用的这些工具的价值榨到最大,你必须把你自己从瓶颈的位置上挪开。你不能一直在那儿等着去提示下一步。你得把自己挪到外面去。你必须把事情安排成它们完全自主。然后越是——你懂的,你怎么才能把你的 token 吞吐量最大化、同时又不在循环里?这才是目标。所以我大概提到,现在这个游戏的核心是提高你的杠杆率。我只是偶尔放进去很少的几个 token,然后大量的事情就替我自动发生了。所以 auto research——我发了那条推,我觉得人们挺喜欢的,但他们也许还没有把它的推论想透,而对我来说 auto research 就是那个推论的一个例子。它的意思是,我不想当那个在循环里的研究者——盯着结果看之类的。我那是在拖累系统。所以问题就是,我怎么去重构所有这些抽象,让我不必——我只需要安排一次然后按下"开始"。这个游戏的核心是:你怎么才能让更多的 agent 在你不介入的情况下、更长时间地运行,替你做事?而 auto research 就是——是啊,这是一个目标,这是一个指标,这是你能做和不能做的边界。然后开跑。然后,是啊,它成功了。
[17:43]
at its effectiveness. Yeah, I I didn't expect, uh, it to work because so I have the project data chat, um, and fundamentally like I think a lot of people are very confused with my obsession for like training GPT-2 models and so on. But for me, uh, training GPT models and so on is just a little harness, a little playground for training LLMs. And fundamentally what I'm more interested in is like this idea of recursive self-improvement and to what extent you can actually have LLMs improving LLMs because I think all the frontier labs this is like the thing Mhm. uh, for obvious reasons and they're all trying to recursively self-improve roughly speaking. And so for me this is kind of like, um, a little playpen of that. Um, and I guess I like tuned Nan Chat already quite a bit by hand in the good old fashion way that I'm used to. Like I'm a researcher. I've done this for like, you know, two decades. I have some amount of like What is the opposite of hubris? Uh, yeah. [laughter] Earned confidence? Okay. I have like two decades of like, "Oh, I've trained this model like thousands of times. I've like, um, so I've done a bunch of experiments. I've done hyperparameter tuning. I've done all the things I'm very used to and I've done for two decades. Yeah. And I've gotten to a certain point and I thought it was like fairly well tuned and then I let auto research go for like overnight and it came back with like tunings that I didn't see. Mhm. And yeah, I did forget like the weight decay on the value embeddings and my Adam betas were not sufficiently tuned and these things just jointly interact. So like once you tune one thing the other things have to potentially change too. You know, I shouldn't be a bottleneck. I shouldn't be running these hyperparameter optimizations. I shouldn't be looking at the results. There's objective criteria in this case. Uh, so you just let you just have to arrange it so that it can just go forever. So that's a single sort of version of auto research of like a single loop trying to improve. And I was surprised that it, um, it found these things that I you know, the repo was already fairly well tuned and still found something. And that's just a single it's a single loop. Like these frontier labs they have GPU clusters of tens of thousands of them. And so it's very easy to imagine how you would basically get a lot of this automation on, um, smaller models. And fundamentally everything around like frontier level intelligence is about extrapolation and scaling loss. And so you basically do a ton of the exploration on the smaller models and then you try to, um, extrapolate out. So you're saying our research efforts are going to get more efficient. Like we're going to have better direction for when we scale as well if we can do this experimentation better.
它的效果。是啊,我没料到它会成功,因为——我有个项目叫 data chat,本质上——我觉得很多人对我执着于训练 GPT-2 模型之类的事情感到非常困惑。但对我来说,训练 GPT 模型之类的只是一个小框架、一个用来训练 LLM 的小游乐场。我真正更感兴趣的,本质上是"递归式自我改进"这个想法,以及你究竟能在多大程度上真正做到让 LLM 去改进 LLM。因为我觉得所有前沿实验室——这就是那个核心命题,嗯,出于显而易见的原因,他们大致上都在试图实现递归式自我改进。所以对我来说这就有点像那件事的一个小小的游乐场。我猜我已经用我习惯的那种老派方式、亲手把 Nan Chat 调得相当不错了。我是个研究者。我做这个做了大概——你懂的——二十年了。我有一定程度的——"狂妄自大"的反义词是什么来着?嗯,是啊。[笑声] "挣来的自信"?好吧。我有大概二十年的那种经验,"哦,这个模型我已经训练过成千上万次了",我做过一大堆实验、做过超参调优、做过所有这些我非常熟悉、做了二十年的事情。是啊。我把它调到了某个程度,我觉得它已经调得相当不错了,然后我让 auto research 跑了一个通宵,它回来给了我一些我自己没看到的调参。嗯。是啊,我确实忘了 value embedding 上的 weight decay,而且我的 Adam betas 调得不够充分,这些东西是会联合相互作用的。所以你一旦调了一个东西,别的东西也可能得跟着改。你懂的,我不该成为瓶颈。我不该亲自去跑这些超参优化。我不该去盯着结果看。在这种情况下是有客观判据的。所以你就——你只需要把它安排好,让它能一直跑下去。所以这是 auto research 的一种单一版本,就是一个单一的、努力做改进的循环。我很惊讶它居然——它居然找出了那些东西,你懂的,那个仓库本来已经调得相当不错了,它还是找到了点东西。而这只是一个单一的——它就是一个单一的循环。而这些前沿实验室,他们有几万台 GPU 的集群。所以很容易想象,你基本上就能在更小的模型上得到大量这样的自动化。而本质上,所有围绕"前沿级智能"的东西都是关于外推和 scaling law 的。所以你基本上是在更小的模型上做大量的探索,然后再试着往外推。所以你的意思是,我们的研究努力会变得更高效。比如说我们会对"何时该扩大规模"有更好的方向感,前提是我们能把这种实验做得更好。
[19:50]
Yeah, I would say that like the most interesting project and probably what the frontier labs are working on is uh, Mhm. Yeah. you know, you experiment on the smaller models. You try to make it as autonomous as possible. Remove researchers
是啊,我会说,最有意思的项目、也很可能是前沿实验室正在做的事情,就是——嗯,是啊,你在更小的模型上做实验。你尽量把它做得越自主越好。把研究者……
[19:59]
[laughter]
[笑声]
[20:00]
from the loop. They have way too much What is the What is the opposite of too much confidence? Yeah, yeah, they don't know. They shouldn't be touching any of this really. And so you have to like rewrite the whole thing because right now, I mean certainly they can contribute ideas. But okay, they shouldn't actually be enacting these ideas. There is a queue of ideas and there's maybe an automated scientist that comes up with ideas based on all the archive papers and GitHub repos and it funnels ideas in or researchers can contribute ideas, but it's a single queue and there is workers that pull items and they try them out. And whatever works just gets sort of put on the feature branch and maybe some people like monitor the feature branch and merge to the main branch sometimes. So yeah, just removing humans from all the processes and automating as much as possible and getting high token tokens per second throughputs and it does require rethinking of all the abstractions and everything has to be reshuffled. So yeah, I think it's very exciting. If we take one more recursive step here, when is the model going to write a better program MD than you? Yeah. Also program MD is like
……从循环里移出去。他们的——"过度自信"的反义词是什么来着?是啊是啊,他们不知道。他们其实根本就不该碰这些东西。所以你得把整套东西重写一遍,因为现在——我是说他们当然可以贡献想法。但是,好吧,他们其实不该亲自去执行这些想法。有一个想法的队列,也许有一个自动化的科学家,它基于所有的 arxiv 论文和 GitHub 仓库提出想法,把想法灌进队列里,或者研究者也可以贡献想法,但这是一个单一的队列,然后有一些 worker 会从里面取出条目去尝试。凡是有效的就会被放到 feature 分支上,也许有一些人会监控这个 feature 分支,时不时合并到 main 分支。所以是啊,就是把人类从所有流程里移出去,尽可能多地自动化,把每秒 token 吞吐量做到很高,而这确实需要重新思考所有的抽象,一切都得重新洗牌。所以是啊,我觉得这非常令人兴奋。如果我们在这里再往递归走一步——模型什么时候会写出一个比你更好的 program.md?是啊。而且 program.md 本身就是个……
[21:03]
loop. Yeah, exactly.
……循环。是啊,没错。
[21:05]
Yeah. So program MD is my crappy attempt at describing like how the auto researcher should work. Like oh, do this then do that and that and then try these kinds of ideas and then here's maybe some ideas like look at architecture, look at optimizer, etc. But I just came up with with this in markdown, right?
是啊。所以 program.md 是我用来描述"auto researcher 应该怎么工作"的一个很糙的尝试。比如说,哦,先做这个再做那个再做那个,然后试试这几类想法,然后这里也许有一些想法,比如看看架构、看看 optimizer 等等。但我就是用 markdown 把这个东西写出来了,对吧?
[21:19]
Mhm. And so yeah, exactly. You want some kind of an auto research loop maybe that looks for You can imagine that different program that MDs would would give you different progress. So you basically every research organization is described by program MD. A research organization is a set of markdown files that describe all the roles and how the whole thing connects. And you can imagine having a better research organization. So maybe they do fewer stand-ups in the morning because they're useless. And this is all just code, right? And so you can So one organization can have fewer stand-ups, one organization can have more. One organization can be very risk-taking, one organization can be less. As you can definitely imagine that you have multiple research orgs and then they all have code. And once you have code, then you can imagine tuning the code. So 100% there's like the metal layer of it. Uh Did you see my text about my contest idea? My contest idea was like let people write different program MDs, right? And and so for same hardware, where do you get most improvement?
嗯。所以是啊,没错。你会想要某种 auto research 循环,也许它会去寻找——你可以想象,不同的 program.md 会给你带来不同的进展。所以本质上每一个研究组织都是由一个 program.md 来描述的。一个研究组织就是一组 markdown 文件,描述所有的角色以及整个系统是怎么连接起来的。然后你可以想象拥有一个更好的研究组织。所以也许他们早上少开几次站会,因为站会没用。而这一切都只是代码,对吧?所以你可以——一个组织可以少开站会,一个组织可以多开。一个组织可以非常敢于冒险,一个组织可以保守一些。你完全可以想象,你有多个研究组织,然后它们都有代码。而一旦你有了代码,你就可以想象去调优这些代码。所以百分之百,这里面有一个"元层"。你看到我那条关于比赛点子的消息了吗?我的比赛点子是这样的——让人们去写不同的 program.md,对吧?然后在同样的硬件上,看谁能拿到最大的提升。
[22:22]
Oh, I see. And then you can take all that data and then give it to the model and say write a better program MD.
哦,我明白了。然后你可以把那些数据全部拿过来,喂给模型,跟它说:写一个更好的 program.md。
[22:26]
Yes, yes. Yeah, exactly.
对,对。没错,就是这样。
[22:28]
We're going to get something better. Like there's no way we don't, right?
我们肯定会得到更好的东西。不可能得不到吧,对吧?
[22:30]
100% look at where the improvements came from and like can I change the program MD such that more of these kinds of things would be done or like things that didn't work except you can 100% imagine doing that. So I think this is a great idea, but it's like you know, I think like you can sort of go one step at a time where you sort of have one process and then second process and then the next process and these are all layers of an onion. Like the LLM sort of part is now taken for granted. The agent part is now taken for granted. Now the claw-like entities are taken for granted and now you can have multiple of them and now you can have instructions to them and now you can have optimization over the instructions and it's just like a little too much, you know, but I mean this is why it gets to the psychosis is that this is like infinite and everything is scale issue and that's why I feel like Yeah, that's just coming back to This is why it's so insane. Okay, well, if [laughter] we're we're just trying to like diagnose the current moment and what is a relevant skill right now, what do you like what do you think is the implication that this that this is the loop we should be trying to achieve in different areas and then it works, right? Like you know, remove create the metric or create the ability for agents to continue working on it without you. Do we still have performance engineering? Like what Yeah, I mean so there's a few caveats that I would put on top of the LLM psychosis. So number one, this is extremely well suited to anything that has objective metrics that are easy to evaluate. So for example, like writing kernels for more efficient CUDA, you know, code for various parts of the model, etc. are a perfect fit because you have inefficient code and then you want efficient code that has the exact same behavior but it's much faster. Perfect fit. So a lot of things like like are perfect fit for auto research, but many things will not be. And so they it's just if you can't evaluate then you can't auto research it, right? So that's like caveat number one. And then maybe caveat number two I would say is you know, we're we're kind of talking about the next steps and we kind of see what the next steps are, but fundamentally the the whole thing still doesn't it still kind of like bursting at the seams a little bit and there's cracks and it doesn't fully work and if you kind of try to go too far ahead, the whole thing is actually net not useful if that makes sense. Because these models like still are not, you know, they've improved a lot, but they're still are like rough around the edges is maybe the way I would describe it. I simultaneously feel like I'm talking to an extremely brilliant PhD student who's been like a systems programmer for their entire life and a 10-year-old. And it's so weird because humans like there's like I feel like they're a lot more coupled like you have to you know, um Yes, you wouldn't you wouldn't encounter that combination.
百分之百会。你去看看进步是从哪儿来的,然后想:我能不能改一改 program.md,让这类有效的事情多发生一些,或者把那些没奏效的去掉。你完全可以想象去做这件事。所以我觉得这是个很棒的主意。但你知道,我觉得你可以一步一步来——你先有一个流程,再有第二个流程,再有下一个流程,这些就像洋葱的一层层。LLM 那部分现在已经被视为理所当然了;agent 那部分现在也被视为理所当然了;现在 Claude 那样的实体也被视为理所当然了;现在你可以同时跑好几个,可以给它们下指令,可以对这些指令做优化——这就有点太多了,你懂吗。但我想说,这正是为什么会陷入那种精神错乱的状态——因为这东西是无限的,一切都是规模问题,这也是为什么我觉得……对,又回到这一点了,这就是为什么它这么疯狂。好吧,那……[笑声] 我们只是想诊断一下当下这个时刻,以及现在什么才是相关的技能。你觉得这意味着什么?这是不是我们应该在不同领域努力实现的那个循环,然后它就奏效了,对吧?你知道的——去掉、去创造那个指标,或者创造让 agent 能在你不在场的情况下继续干活的能力。我们还需要性能工程师吗?就像……对,我是说,关于这种 LLM 精神错乱,我想加几个限定条件。第一,这种做法极其适合任何有客观、容易评估的指标的事情。比如说,给模型的各个部分写更高效的 CUDA kernel 之类的代码,就是绝佳的契合点——你有一段低效的代码,然后你想要一段行为完全一样但快得多的高效代码。完美契合。所以很多事情都非常适合 auto research,但很多事情就不适合。如果你没法评估,你就没法对它做 auto research,对吧?这是第一个限定条件。第二个限定条件大概是这样:我们现在在聊接下来的步骤,我们也大概看到了接下来的步骤是什么,但从根本上说,整个东西还是有点……还是有点接缝处快撑爆的感觉,有裂缝,没法完全跑通。如果你想往前冲得太远,整个东西其实净效果上是没用的,如果你懂我意思的话。因为这些模型……还是不行,你知道,它们进步了很多,但还是有点毛糙——也许我会这么形容。我同时感觉自己像是在跟一个极其聪明的、一辈子都在做系统编程的博士生说话,又像是在跟一个十岁小孩说话。这太怪了,因为人类……我感觉人类身上这两者要耦合得多,你必须……你知道,嗯……对,你不会碰到那种组合。
[24:54]
This jaggedness is really strange and humans have a lot less of that kind of jaggedness, although they definitely have some.
这种参差不齐真的很奇怪,人类身上这种参差要少得多,虽然人类肯定也有一些。
[24:59]
[laughter]
[笑声]
[25:00]
But humans have a lot more jaggedness. Uh sorry, the agents have a lot more jaggedness where sometimes like you know, I ask for functionality and it like comes back with something that's just like totally wrong and then we get into loops that are totally wrong and then I'm just I get so frustrated with the agents all the time still because you feel the power of it, but you also there's still like it does not say statistical things once in a while for me as well. I get very annoyed [clears throat] when I feel like the agent wasted a lot of compute on something it should have recognized was an obvious problem. Yeah. I think like some of the bigger things is like maybe what's under underneath it if I could hypothesize is fundamentally these models are trained via reinforcement learning. So they're actually struggling with the exact same thing we just talked about which is the labs can improve the models in anything that is verifiable or that [clears throat] has rewards. So did you write the program correctly and does it you do you the unit tests check out? Yes or no. But some of the things where they're struggling is like for example, I think they have a tough time with like nuance of maybe what I what I had in mind or what I intended and when to ask clarifying questions. Um or like what I Yeah, it's just um anything that feels softer is like worse. And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders. Like maybe another way to put it is if you go to if today if you go to like state-of-the-art model, ChatGPT and you ask it tell me a joke, um do you know what joke you're going to get? There's the joke. The joke? I do feel I I I can't tell you like the you know, standard form of it, but I do feel like ChatGPT has like three jokes.
但人类的参差要少得多。呃,抱歉,我是说 agent 的参差要多得多——有时候你知道,我让它实现某个功能,它返回来的东西完全是错的,然后我们就陷进了完全错误的循环里,我到现在还是经常被 agent 气得不行。因为你能感受到它的强大,但同时……它还是会时不时给我冒出一些不靠谱的东西。当我觉得 agent 在一个本该一眼看出是明显问题的事情上浪费了一大堆算力时,我会非常恼火 [清嗓子]。对。我觉得,如果让我做个假设,更深层的一些原因可能是:这些模型本质上是通过强化学习训练出来的。所以它们其实正在跟我们刚才聊的同一件事较劲——实验室能在任何可验证的、或者有奖励信号的东西上 [清嗓子] 改进模型。比如:你这个程序写对了吗?单元测试通过了吗?通过还是没通过。但它们吃力的一些地方是,比如说,我觉得它们很难拿捏那种微妙之处——我心里到底想要的是什么、我的意图是什么,以及什么时候该问澄清性的问题。呃,或者像我……对,就是说,任何感觉上更"软"的东西它就做得更差。所以你基本上是这样:你要么在轨道上,是超级智能那套电路的一部分;要么你不在轨道上,跑出了可验证的领域,然后一切就突然变得漫无目的、东游西荡。换个说法吧——如果你今天去用最先进的模型,ChatGPT,你跟它说"给我讲个笑话",你知道你会得到什么笑话吗?就那个笑话。那个笑话?我确实觉得……我说不出它标准的那个版本,但我确实觉得 ChatGPT 大概就那么三个笑话。
[26:34]
Yeah, yeah. So the the joke that apparently all the LLMs like love the most is why do scientists not trust atoms? Okay. Because they make everything up. Okay.
对,对。所有 LLM 显然最爱的那个笑话是:科学家为什么不信任原子?嗯。因为原子能编造(make up)一切。嗯。
[26:44]
They make everything up. So this is still
它们编造一切。所以这个还是……
[26:46]
emerge? So this is the joke you would get like three or four years ago and this is the joke you still get today. Okay.
冒出来?所以这是你三四年前会得到的那个笑话,今天你还是会得到同一个笑话。嗯。
[26:52]
So even though the models have improved tremendously and if you give them an agentic task, they will just go for hours and move mountains for you. And then you ask for like a joke and it has a stupid joke. It's crappy joke from five years ago and it's because it's outside of the it's outside of the RL. It's outside of the reinforcement learning. It's outside of what's being improved. It's like and it's part of the jaggedness of like shouldn't you expect models as they get better to also have like better jokes or more diversity of them or it's just it's not being optimized and stuck. Do you think that that implies that we are not seeing like generalization in the sense of like broader intelligence of joke smartness being attached to code smartness? Yeah, I think there's some decoupling where some things are verifiable and some things are not and some things are optimized for arbitrarily by the labs depending on like what data went in and some things are not and um and
所以哪怕模型已经进步得这么厉害——你给它一个 agentic 的任务,它能连续干上好几个小时,为你移山倒海——可你一让它讲个笑话,它就讲一个蠢笑话。一个五年前那种烂笑话。这是因为它在 RL 之外,在强化学习之外,在被改进的范围之外。这就是参差不齐的一部分——你难道不该期待模型变得更好的同时,笑话也变得更好、或者笑话更多样化吗?可它就是没被优化,卡在那儿了。你觉得这是不是意味着,我们并没有看到那种泛化——也就是说,更广义上的智能、讲笑话的聪明,并没有跟写代码的聪明绑在一起?对,我觉得确实存在某种解耦——有些东西可验证,有些不可验证;有些东西被实验室随意地优化了,取决于喂进去的是什么数据,有些则没有。嗯,而且……
[27:46]
But I mean the the premise there's a you know, premise from some research groups that if you're smarter at code generation or in these verifiable fields, you should be better at everything. And like the the joke situation suggests that that's not happening at all. Okay.
但我是说,那个前提——你知道,某些研究团队有一个前提:如果你在代码生成上更聪明、在这些可验证的领域里更聪明,那你应该在所有事情上都更厉害。而笑话这个情况说明,这种事根本就没发生。嗯。
[28:01]
Yeah, I don't think that's happening. I think I think maybe we're seeing like a little bit of that, but not like a satisfying amount.
对,我不认为那种事正在发生。我觉得……也许我们能看到一点点那种迹象,但远没有到让人满意的程度。
[28:06]
Yeah, that jaggedness exists in humans. You [laughter] can be very very good at math and still tell really bad jokes.
对,这种参差不齐在人类身上也存在。你 [笑声] 可以数学非常非常好,但讲的笑话还是烂得要命。
[28:13]
Yeah, that's true. Yeah, but it just it still means that we're not getting like the story is that we're getting a lot of the intelligence and capabilities in all the domains of society like for free as we get better and better models and that's not like exactly fundamentally what's going on and there's some blind spots and some things are not being optimized for and this is all clustered up in these neural net opaque models, right? So you're either on rails of what it was trained for and everything is like you're going at speed of light or you're not. And so it's the jaggedness. So um So that's why I think like even though the the progression is obvious what should happen, you can't let it fully go there yet because it doesn't fully work or it's a scale issue and we just haven't like figured out how to use it. So you know, it's hard to tell. Can I ask a somewhat blasphemous question which is like if this jaggedness is persisting and it's all rolled up in a at least monolithic interface, right? But you know, single model. Does that make sense or do you should should it be unbundled into things that are can be optimized and improved against different domains of intelligence? Like unbundling the models into multiple experts in different areas, etc. More directly. Yeah. Um Instead of just MOE that we have no exposure to because that can be like confusing as a user from the outside which is like why is it so good at this, but not at this other thing? Yeah, I think currently my impression is the labs are trying to have a single sort of like monoculture of a model that is arbitrarily intelligent in all these different domains and they just stuff it into the parameters. I do think that we will we I do think we should expect more speciation in the intelligences. Um like, you know, the animal kingdom is extremely diverse in the brains that exist and there's lots of different niches of of nature and some animals have overdeveloped visual cortex or other part kind of parts and I think we we should be able to see more speciation and um you don't need like this oracle that knows everything. You can speciate it and then you put it on a specific task and we should be seeing some of that because you should be able to have like much smaller models that still have the cognitive core like they're still competent but then they specialize and then um and then they they can become more efficient in terms of latency or throughput on specific tasks that you really care about. Like if you're a mathematician working in Lean, I saw for example there's a few releases that really like target that as a domain. Um uh so there's a probably going to be a few examples like that where the unbundling kind of makes sense. One question I have is whether or not the capacity constraint on available compute infrastructure Mhm. drives more of this because efficiency Yeah. actually matters more. Yeah. Your if you financing aside, though financing's involved in all of this. If you have access to full compute for anything you do like even one single model, right? But if you actually feel pressure where you're like I can't serve
对,没错。对,但这仍然意味着我们并没有得到那种……人们讲的故事是:随着模型越来越好,社会各个领域的智能和能力我们都能免费得到。可从根本上说,实际发生的并不完全是这么回事——有一些盲点,有些东西没被优化,而这一切全都团成一坨,挤在这些不透明的神经网络模型里,对吧?所以你要么在它被训练过的那个轨道上,一切都以光速狂奔;要么你不在。这就是参差不齐。嗯。所以这就是为什么我觉得,尽管该发生什么是显而易见的,你现在还不能完全放手让它往那个方向去,因为它还没完全跑通,或者说这是个规模问题,我们只是还没搞清楚怎么用它。所以你知道,很难说。我能问一个有点大不敬的问题吗?就是:如果这种参差不齐一直存在,而它又全都团在一个——至少是一个单一的接口里,对吧?你知道,单个模型。这说得通吗?还是说,它应该被拆开,拆成那些可以针对不同智能领域分别优化、分别改进的东西?也就是把模型解绑成多个在不同领域的专家之类的——更直接地。对。嗯——而不只是我们从外部完全看不见的那种 MoE,因为那个从用户的外部视角看会让人很困惑:为什么它这件事这么强、那件事却不行?对,我觉得目前我的印象是,实验室在试图打造一个单一的、单一栽培式的模型,让它在所有这些不同领域里都任意地聪明,然后把这一切全塞进参数里。但我确实觉得我们应该预期智能体之间会出现更多的物种分化。嗯,你知道,动物界里存在的大脑极其多样,自然界有很多不同的生态位,有些动物有过度发达的视觉皮层或者别的某些部分。我觉得我们应该能看到更多的物种分化。而且,你不需要那种无所不知的神谕。你可以让它分化,然后把它放到一个特定任务上。我们应该能看到一些这样的迹象,因为你应该能拥有那种小得多的模型——它们仍然有那个认知内核,仍然很能干,但它们专精了,然后在你真正在意的特定任务上,它们在延迟或吞吐量方面能变得更高效。比如你是个用 Lean 工作的数学家,我就见过有几个发布版本确实是把这个当作目标领域来做的。呃,所以大概会有几个这样的例子,在那些情形里解绑是说得通的。我有一个问题:可用算力基础设施的容量约束,会不会推动更多这种分化——因为效率……嗯……其实变得更重要了。对。先把融资放一边——虽然融资跟这一切都有关系。如果你想做什么就有充足的算力,哪怕只用单个模型,对吧?但如果你真的感受到压力,那种"我没法服务……"的压力……
[30:59]
Mhm. um model of massive size for every use case.
嗯。……没法为每一个使用场景都上一个巨大尺寸的模型。
[31:03]
Mhm. Like do you think that leads to any speciation? Does that question make sense to you? The question makes sense and I guess like what I'm what I'm what I what I'm struggling with is I don't think we've seen too much speciation just yet, right? No. Uh we're seeing a monoculture of models. Yeah. So um And there's like clearly pressure for like make a good code model, put it back in the main, merge again. Yeah.
嗯。那你觉得这会不会导致某种物种分化?这个问题你听得懂吗?这个问题听得懂。我猜,我现在纠结的地方是——我觉得我们目前还没看到太多物种分化,对吧?没错。呃,我们看到的是一种单一栽培式的模型。对。所以呢,嗯……而且显然存在那种压力:做一个好的代码模型,再把它合回主干,再次合并。对。
[31:23]
Um even though there already is pressure on the models. Mhm. I guess perhaps I I feel like there's a lot of very short-term supply crunch and like maybe that causes more speciation now. Yeah, I think fundamentally like the the the labs are serving a model and they don't really know what the end user is going to be asking about. So maybe that's like some part of it because they kind of have to multitask over all the possible things they could be asked. But I think if you're coming to a business and maybe partnering on some specific problems you care about then maybe you would see that there. Um or there would be some very high-value applications that are like more niche. Um But but I think right now they're kind of like going after the totality of what's available. I don't think that the science of manipulating the brains is like fully developed yet partly. What do you mean manipulating? So like so fine-tuning without losing capabilities as an example. And I we don't have these primitives for actually like working with the intelligences in ways other than just context windows. Our context windows kind of just just work and it's very cheap to manipulate etc. And this is how we're getting some of the customization etc. Uh but I think if it was I think it's a it's a bit more of a developing science of how you like more deeply adjust the models, how you have continual learning maybe or how you um how you fine-tune in a certain area, how you get better in a certain area or like how you actually touch the weights not just the context windows. And so it's a lot more tricky I would say to touch the weights than just the context windows uh because you're actually fundamentally changing the full model and potentially its intelligence. And so um so maybe it's just like not a fully developed science if that makes sense of speciation. And it also has to be like cheap enough Yeah. for that speciation to be worthwhile in these given
嗯,尽管模型上已经存在压力了。嗯。我猜,也许吧,我感觉现在有很多非常短期的供给紧张,而这或许会在当下催生更多的物种分化。对,我觉得从根本上说,实验室在服务一个模型,他们其实并不知道终端用户会问什么。所以这或许是原因的一部分——因为他们某种程度上必须在所有可能被问到的事情上做多任务处理。但我觉得,如果你来找一家企业,针对你在意的某些具体问题做合作,那也许在那种情况下你就会看到分化。嗯,或者会有一些价值非常高、比较小众的应用。嗯,但我觉得现在他们某种程度上是在追求"可获得的全部"。我觉得操控这些"大脑"的科学还没完全成熟,部分原因……你说的"操控"是什么意思?就是说,比如在不丢失能力的前提下做微调(fine-tuning)。我们还没有那些真正的原语,能让我们用除了 context window 之外的方式去跟这些智能体打交道。我们的 context window 某种程度上就是直接好用,操控它也很便宜,等等。这就是我们现在获得一部分定制化的方式。呃,但我觉得,关于你如何更深层地调整模型、你如何做持续学习、你如何在某个领域做微调、你如何在某个领域变得更好、或者你如何真正去碰权重而不只是 context window——这门科学还更不成熟一些。我会说,去碰权重比只碰 context window 要棘手得多,因为你实际上是在从根本上改变整个模型,可能还改变它的智能。所以嗯,所以也许物种分化这门科学就是还没完全成熟,如果你懂我意思的话。而且它还得足够便宜……对……物种分化才值得在这些给定的……
[32:57]
contexts. Can I ask a question about like an extension to auto research that you described in terms of open ground? You say okay, well, you know, we have this thing. Um we need more collaboration surface around it essentially for people to contribute to research overall. Can you talk about that?
……场景里去做。我能问一个问题吗,关于你描述的 auto research 的一个延伸,从"开放场地"的角度。你说,好吧,你知道,我们有了这个东西。嗯,我们本质上需要围绕它建立更多的协作界面,让大家都能为整体研究做贡献。你能聊聊这个吗?
[33:15]
Yeah, so we talked about auto research has a single thread of like I'm going to try stuff in a loop but fundamentally the parallelization of this is like the interesting component. And I guess I was trying to like play around with a few ideas but I don't have anything that like clicks as simply as like I don't have something I'm like super happy with just yet but it's something I'm like working on the side when I'm not working on my claw. Um so I think like one issue is if you have a bunch of nodes of parallelization available to then it's very easy to just have multiple auto researchers talking through a a common system or something like that. What I was more interested in is how you can have an untrusted pool of workers out there on the internet. Mhm. So for example in auto research you're just trying to find um the piece of code that trains a model to a very low validation loss. If anyone gives you a candidate commit, it's very easy to verify that that commit is correct is good. Like they someone could claim from the internet that this piece of code will optimize much better and give you much better performance. You could just check. Yeah. But probably a lot of work goes into that checking. But fundamentally they could lie and etc. So you're basically dealing with a similar kind of it's almost actually like looks a little bit like my my designs that incorporate an untrusted pool of workers actually look a little bit more like a blockchain a little bit uh because instead of blocks you have commits and these commits can build on each other and they contain like changes to the code as you're improving it. Um and uh the proof of work is basically doing tons of experimentation to find the commits that work. Um and that's hard and then the reward is just being on the leaderboard right now. There's no monetary reward whatsoever. Uh but I don't want to push the analogy too far but it fundamentally has this issue where you a huge amount of search goes into it but it's very cheap to verify that a candidate solution is indeed good because you can just train a single you know, someone had to try 10,000 ideas but you just have to check that the thing that they produced actually works because the 99,000 of them didn't work, you know? Um and so basically long story short is like you have to come up with a system where an untrusted pool of workers can collaborate with a trusted pool of workers that do the verification. And the whole thing is kind of like asynchronous and works and and so on and it's it's like safe from a security perspective because if anyone sends you arbitrary code and you're going to run it, that is very sketchy and dodgy. So um but fundamentally it should be totally possible. So you're familiar with projects like SETI@home and Folding@home. All of these problems have a similar kind of setup. So Folding@home you're folding a protein and it's very hard to find a configuration that is low energy. But if someone finds a configuration that they value to be low energy, that's perfect. You can just use it. You can easily verify it. So a lot of things have this property that you know, very expensive to come up with but very cheap to verify. And so in all those cases things like Folding@home or SETI@home or auto research at home will be good fits. And so um long story short a swarm of agents on the internet could collaborate to improve LLMs and could potentially even like run circles around frontier labs. Like who knows, you know? Um yeah, like maybe that's even possible. Like frontier labs have a huge amount of trusted compute but the earth is much bigger and has huge amount of untrusted compute. But if you put systems in check systems in place that you know, deal with this then maybe it is possible that the swarm out there could could come up with with better with better solutions. And people kind of like contribute cycles um to to a thing that they care about. And so sorry to so the last thought is uh lots of companies or whatnot they could maybe have like their own things that they care about and you if you have compute capacity you could contribute to different kind of auto research tracks. Like maybe you care about certain you know, like you care about like cancer or something like that of certain type. You don't have to just donate money to an institution. You actually could like purchase compute and then you could join the auto research swarm for that project, you know? Uh so if everything is rebundled into auto researchers then compute becomes the thing that you're contributing to the pool. Yeah. That's very inspiring and it's also interesting. Like I don't I don't know how far this goes but it is interesting that at least some audience of people you know, here in Silicon Valley or lining up at you know, retail stores in China have discovered that like having access to personal compute is interesting again.
对,所以我们聊过 auto research 是单线程的——就是我会在一个循环里不停尝试各种东西。但从根本上说,把这件事并行化才是有意思的部分。我之前一直在试着琢磨几个想法,但还没有什么能像……我还没找到一个让我特别满意的、那种简单得"咔哒"一下就成立的东西。不过这是我在不忙我那个 Claude 的时候,在业余时间一直在搞的东西。所以我觉得,一个问题是:如果你手头有一堆可用的并行节点,那要让多个 auto researcher 通过一个公共系统之类的东西互相交流,是很容易的。但我更感兴趣的是,你怎么能利用互联网上一个不可信的工人池。嗯。比如说,在 auto research 里,你只是想找到那段能把模型训练到极低验证损失的代码。如果有人给你一个候选 commit,要验证那个 commit 是不是正确、是不是好,是很容易的。比如有人可能在互联网上声称,这段代码能优化得好得多、给你带来好得多的性能。你只要去检验一下就行。对。不过那个检验过程大概要花不少功夫。但从根本上说,他们可能会撒谎,等等。所以你基本上面对的是一种类似的……其实,我那些纳入了不可信工人池的设计,看起来还真有点像区块链——有点像——因为不是区块(block),你这里是 commit,而这些 commit 可以一个叠一个地往上建,它们包含的是你在改进代码过程中对代码的修改。嗯,而工作量证明(proof of work)基本上就是做海量的实验,去找到那些奏效的 commit。嗯,那很难,然后奖励就只是上排行榜——现在压根没有任何金钱奖励。呃,我不想把这个类比推得太远,但它从根本上有这个特性:你要往里投入海量的搜索,但验证一个候选方案是不是真的好却非常便宜——因为你只要训练一个……你知道,有人可能得试一万个点子,但你只需要检验他们产出的那个东西是不是真的奏效,因为那一万里有九千九百多个都没奏效,你懂吗?嗯,所以长话短说就是:你得设计出一个系统,让一个不可信的工人池能跟一个负责做验证的可信工人池协作。整个东西某种程度上是异步的,并且能跑通,等等,而且从安全角度看它要是安全的——因为如果随便什么人给你发来任意代码而你又要去运行它,那是非常可疑、非常危险的。所以嗯,但从根本上说这应该完全是可行的。你应该熟悉 SETI@home、Folding@home 这类项目。所有这些问题都有一种类似的结构。Folding@home 里你是在折叠蛋白质,要找到一个低能量的构型非常难。但如果有人找到了一个他们认定是低能量的构型,那就完美了,你直接拿来用就行。你可以轻松验证它。所以很多事情都有这个性质:想出来非常昂贵,但验证非常便宜。在所有这些情形里——像 Folding@home、SETI@home,或者"auto research@home"——都会是很好的契合。所以嗯,长话短说,互联网上一群 agent 组成的蜂群可以协作来改进 LLM,甚至有可能把前沿实验室耍得团团转。谁知道呢,你知道?嗯,对,也许那甚至是可能的。前沿实验室有海量的可信算力,但地球大得多,有海量的不可信算力。但如果你建立起那些系统、那些制衡机制来应对这个问题,那也许外面那个蜂群真的有可能想出更好的、更好的解决方案。而人们某种程度上是在为他们在意的某个东西贡献算力周期。所以——抱歉——最后一个想法是:呃,很多公司之类的,也许可以有他们自己在意的东西,而如果你有算力容量,你就可以为不同的 auto research 赛道做贡献。比如也许你在意某种……你知道,比如你在意某种特定类型的癌症之类的。你不必只是给某个机构捐钱。你其实可以去买算力,然后加入那个项目的 auto research 蜂群,你懂吗?呃,所以如果一切都被重新打包成 auto researcher,那么算力就成了你往这个池子里贡献的东西。对。这非常鼓舞人心,也很有意思。我不知道这能走多远,但有意思的是——至少有一部分人,你知道,无论是在硅谷这里,还是在中国零售店外排队的那些人——已经发现,拥有个人算力的访问权又重新变得有意思了。
[37:20]
Yeah. Right? So maybe they're really motivated to do that for their claws and then they can contribute to auto research.
对。对吧?所以也许他们真的很有动力为他们自己的 Claude 去搞算力,然后他们就能为 auto research 做贡献。
[37:25]
almost like dollars the thing everyone cares about but is flop the thing that actually everyone cares about in the future? Like is there going to be like a flipening almost of like what's the thing that you care about? Like right now for example it's really hard to get compute even if you have money. Yeah. So actually it almost seems like the flop is like dominant
……就好像,dollar(美元)是现在人人都在意的东西,但 flop(算力)会不会才是未来人人真正在意的东西?会不会出现一种"翻转时刻"——人们在意的那个东西会发生切换?比如说现在,即便你有钱,也很难拿到算力。对。所以其实看起来 flop 几乎已经是主导性的了……
[37:41]
[laughter]
[笑声]
[37:42]
in a certain sense. Um Yeah, so so maybe that's kind of like that. Kind of like that. Like how much how many flops do you control instead of like what wealth you control? I don't actually think that's true but it's kind of interesting to think about. The last thing you released was like a little bit of jobs data analysis. Is that right? What and might have touched a nerve even though you're just like visualizing some public data.
……在某种意义上。嗯,对,所以也许情况大概就是那样。大概是那样。就是说——你掌控了多少 flop、多少算力,而不是你掌控了多少财富。我其实不认为这是真的,但拿来想一想还挺有意思的。你最近发布的东西是一点点就业数据的分析,对吧?这个……即便你只是把一些公开数据可视化了一下,可能也戳到了某些人的神经。
[38:03]
Yeah. Uh what was you know, what were you curious about? Yeah, I guess I was curious to um I mean everyone is like really it's everyone is really thinking about the impacts of AI on the job market and what's going to look like. So I was just interested to take a look like what does the job market look like? Where are the different roles um and how many people are in different professions? And I was like really just interested to like look through the individual cases and try to think myself about like you know, with these AIs and how they're likely to evolve like are these going to be tools that people are using? Are these going to be displacing tools for these professions? And like what are the current professions and how are they going to change? Are they going to grow or uh adjust to a large extent or like what could be new professions? So it's really just like a way to fuel my own chain of thought about the industry I suppose. Mhm. Um and so yeah, the jobs data basically is just a Bureau of Labor Statistics. They actually have um percent outlook for each profession about how much it's expected to grow over the next I think almost a decade. Uh yeah, I think it's a decade but it was made in 2024. Mhm. We need a lot of health care workers. Yeah. So so they've already made those projections and I'm not sure actually 100% what the methodology was that they they put into their projections. Um I guess I was interested to color things by like if people think that what's like primarily being developed now is this kind of like more digital AI that is kind of like almost like these ghosts or spirit entities that can like interact in the digital world and manipulate a lot of like digital information and they currently don't really have a physical embodiment or presence. And the physical stuff is probably going to go slightly slower because you're manipulating atoms. So flipping flipping bits and and the ability to copy-paste digital information is like makes everything a million times faster than accelerating matter, you know, so Um so energetically, I just think we're going to see a huge amount of activity in the digital space, huge amount of rewriting, huge amount of activity, boiling soup. And I think the we're going to see something that in the digital space goes at the speed of light compared to I think what's going to happen in the physical world to some extent. If it would be the extrapolation. And so I think like
是啊。那个,你之前是对什么感到好奇来着?嗯,我想我当时是很好奇,我是说,现在每个人——真的是每个人——都在认真琢磨 AI 对就业市场的冲击,以及未来会变成什么样。所以我就单纯想看看,就业市场到底长什么样?各种岗位都分布在哪里,不同职业里各有多少人?我那时候真的就是想一个个案例去看,自己琢磨一下,就是说,有了这些 AI、再考虑到它们大概率会怎么演进——这些会变成人们手里在用的工具吗?还是会变成把这些职业取代掉的工具?现在的这些职业是什么样子、它们会怎么变?是会扩张,还是会做出大幅度的调整,又或者会冒出哪些全新的职业?所以这其实就是一种给我自己关于这个行业的思路链条添柴加火的方式吧。嗯。所以那份就业数据基本上就是劳工统计局(Bureau of Labor Statistics)的数据。他们其实给每个职业都做了一个百分比的前景预测,预测它在接下来——我想差不多十年里——预计会增长多少。对,我想是十年,不过这份预测是 2024 年做的。嗯。我们需要大量的医疗护理人员。是啊。所以他们已经做出了这些预测,我其实也不是百分百清楚他们做预测时用的是什么方法论。嗯,我当时是想给这些数据上个色,比如说——如果人们认为现在主要在被开发出来的,是这种更偏数字化的 AI,它有点像幽灵或者精神体一样的存在,能在数字世界里互动、能操纵大量数字信息,而它们现在其实还没有什么物理的实体或者存在感。而物理层面的东西大概会推进得稍微慢一些,因为你要操纵的是原子。所以翻转比特、复制粘贴数字信息的能力,让一切都比加速物质快上一百万倍,对吧。所以从能量层面看,我觉得我们会在数字空间里看到极其庞大的活动量,海量的重写,海量的活动,一锅翻滚的浓汤。我觉得我们会看到数字空间里的某种东西以光速推进,而物理世界里发生的事情某种程度上要慢得多——如果做一个外推的话。所以我觉得就像
[40:01]
[clears throat]
[清嗓子]
[40:01]
there's currently kind of like I think overhang where there can be like a lot of unhubbling almost potentially of like a lot of digital information processing that used to be done by computers and people. And now with AIs there's like a third kind of manipulator of digital information. There's going to be a lot of refactoring in those in those disciplines. Um but the physical world is actually going to be like I think behind that by some amount of time. And so I think what's really fascinating to me is like So that's why I was highlighting the the professions that fundamentally manipulate digital information. This is work you could do from your home, etc. Uh because I feel like those will be like things will change. And it doesn't mean that there's going to be less of those jobs or more of those jobs because it does has to do with like demand elasticity and many other factors. But things will change in these professions because of these new tools and um because of this upgrade to the nervous system of the human superorganism
现在大概存在一种我所说的能力悬置。也就是说,过去那些由计算机和人来完成的大量数字信息处理工作,可能会出现大量的能力解除束缚。而现在有了 AI,就多出了第三种数字信息的操纵者。这些领域里会出现大量的重构。但物理世界其实会比这个落后一段时间。所以让我真正着迷的一点是——这也是为什么我之前特别标注了那些本质上在操纵数字信息的职业。这种工作你在家就能做,等等。因为我觉得这些职业里的东西会发生变化。这并不意味着这类岗位会变少或变多,因为这还涉及需求弹性以及其他很多因素。但这些职业会因为这些新工具而发生变化,也因为人类这个超级有机体的神经系统得到了这么一次升级
[40:50]
[laughter]
[笑声]
[40:50]
if you want to think about it that way. Given the look you had at the data, do you have either any observations or um uh guidance for people facing the job market or thinking about what to study now or what skills to develop? I mean we can all go get like I'm very thankful that I have to like meet people for my job right now.
——如果你愿意这么去理解的话。基于你对这些数据的观察,你有没有什么发现,或者对那些正在面对就业市场、或者在思考现在该学什么、该培养什么技能的人有什么建议?我是说,我们都可以去——我现在挺庆幸我的工作还需要我跟人见面。
[41:07]
Yeah.
是啊。
[41:08]
[laughter]
[笑声]
[41:08]
Yeah, more physical. Yeah. Could you do your work from home though? I could. I think there are relationship parts of it that are hard, but most of it I could. Yeah. I think it's really hard to tell because again like the job market is extremely diverse. I think the answers will probably vary, but uh to a large extent like these tools are extremely new, extremely powerful. And so just being you know, just trying to keep up with it is like the first thing. Um and um yeah, because I think a lot of people kind of like dismiss it or Or they're afraid of it. Or they're afraid of it, etc. As which is totally understandable, of course. Yeah, I think like um it's fundamentally an empowering tool at the moment. Um and these jobs are bundles of tasks. And some of these tasks can go a lot faster. And so people should think of it as primarily a tool that it is right now. Um and I think the long-term future of that is uncertain. Yeah, it's kind of really hard to forecast, to be honest. And like I'm not professionally like doing that really. And I think this is a job of like economists to do properly. You are an engineer though. And like one thing I thought was interesting is that like the demand for engineering jobs is continuing to increase.
对,更偏物理层面的工作。是啊。不过你的工作能在家做吗?能。我觉得里面有一些跟人际关系相关的部分会比较难,但大部分我都能在家做。是啊。我觉得这真的很难讲,因为还是那句话,就业市场极其多样化。我想答案大概会因情况而异,但很大程度上,这些工具极其新、极其强大。所以光是——你知道——光是努力跟上它,这就是第一件要做的事。嗯,因为我觉得很多人有点把它当回事不当回事地打发掉了,或者他们害怕它。或者他们害怕它,等等。这当然完全可以理解。是啊,我觉得它眼下从根本上说是一种赋能的工具。而这些工作都是一捆一捆的任务的组合。其中有些任务可以快很多。所以人们应该主要把它当成它现在就是的那种工具来看待。我觉得这件事的长期未来是不确定的。是啊,说实话挺难预测的。而且我也不是专业在做这个。我觉得这应该是经济学家的活儿,得由他们来正经做。不过你是一名工程师。我觉得有意思的一点是,对工程岗位的需求还在持续增长。
[42:08]
Yeah. Um I I can't tell if that's like a temporary phenomenon. I'm not sure how I feel about it. Yeah, do you know? Yeah, that's like the demand elasticity almost like uh software was scarce, right? And so the reason we don't have more demand for software is just there's its scarcity and it's too expensive.
是啊。嗯,我说不准这是不是一种暂时的现象。我对此的感觉也不太确定。是啊,你知道吗?对,这有点像需求弹性——过去软件是稀缺的,对吧?我们之所以对软件没有更多需求,就是因为它稀缺、太贵了。
[42:22]
So if the barrier comes down, then actually you have the Jevons paradox, which is like you know, you actually the demand for software actually goes up. It's cheaper and there's more More powerful, yeah. The the classical example of this always is the ATMs and the bank tellers uh because there was a lot of like fear that um ATMs and computers basically uh would displace tellers. But what happened is they made like the cost of operation of of a bank branch much cheaper. And so there are more bank branches, so there are more tellers. It's like the canonical example people cite. Uh but basically it's just Jevons paradox. Like something becomes cheaper, so there's a lot of unlocked demand for it. Uh so I do think that that's probably I do have like cautiously optimistic view of this in software engineering where I do think um it does seem to me like the demand for software will be extremely large. Um and it's just become a lot cheaper. And um so I do think that for quite some time um it's very hard to forecast, but it does seem to me like right now at least locally there's going to be more demand for software. Um because software is amazing. It's like you know, digital information processing. You're not forced to use like arbitrary tools that were given to you. They're imperfect in various ways. You're not forced to subscribe to what exists. Code is now ephemeral and it can change and it can be modified. Um and so I think there's going to be a lot of activity in the digital space to like rewire everything in a certain sense. And I think it's going to create a lot of demand for for this kind of stuff. I think long-term um yeah, obviously even with auto research like OpenAI or or you know, Anthropic or these other labs like they're employing what like a thousand something researchers, right?
所以一旦门槛降下来,你其实就会遇上杰文斯悖论——就是说,你对软件的需求实际上反而会上升。它更便宜了,也有更多——更强大了,对。这方面最经典的例子永远是 ATM 机和银行柜员,因为当年大家很担心 ATM 和电脑基本上会把柜员取代掉。但实际发生的是,它们让经营一个银行网点的成本便宜了很多。于是就有了更多的银行网点,也就有了更多的柜员。这是大家常引用的那个典型例子。但本质上它就是杰文斯悖论——某样东西变便宜了,于是它就释放出大量被压抑的需求。所以我确实觉得,在软件工程这件事上,我对此持一种谨慎乐观的看法。在我看来,对软件的需求似乎会极其庞大。而它刚刚变得便宜了很多。所以我确实觉得,在相当长一段时间里——这真的很难预测——但在我看来,至少眼下、局部来看,对软件的需求会更多。因为软件太棒了。它就是——你知道——数字信息处理。你不会被迫去用那些塞给你的、随便什么样的工具,那些工具在各方面都不完美。你不会被迫去将就现成的东西。代码现在是临时的、可变的,可以随时被修改。所以我觉得数字空间里会出现大量活动,某种意义上把一切都重新接线。我觉得这会创造出大量对这类东西的需求。我觉得长期来看,嗯,很显然——即便有了自动研究(auto research),像 OpenAI、Anthropic 或者其他这些实验室,它们雇了多少人,大概一千多名研究员,对吧?
[43:53]
Mhm. These researchers are basically like glorified auto like you know.
嗯。这些研究员基本上就像是被美化了的自动——你懂的。
[43:57]
[laughter]
[笑声]
[43:58]
They're like automating themselves away like actively and this is like the thing they're all trying to do. Yeah. I like I went around um Some of those researchers also fear that feel the psychosis, right? Because they can it's working, right? And and so they're like it's over for me, too. I did spend a bunch of time going around OpenAI and I was like, you guys realize if we're successful like we're all out of job like like this is just going to we're just building automation for Sam or something like that. Like I or the board or I'm not sure, but like uh they're just building all this automation for yeah, the board or the CEO or something like that. And we're all out of our job and maybe contributing on the side. And so yeah, it's kind of like unnerving from that perspective. Is it okay if I ask you Noam's question? Mhm. You know, you could be doing that, right? Auto researching with a lot of compute scale and a bunch of colleagues at one of the frontier [clears throat] labs. Like why not? Well, I was there for a while, right? Like and I did reenter. So to some extent I agree and I think that there are many ways to slice this question. It's very loaded question a little bit. Um I will say that I feel very good about like what people can contribute and their impact outside of the frontier labs, obviously. Not in the industry, but also in like more like ecosystem level roles. Um so your role for example is more like ecosystem level. My role currently is also kind of more on ecosystem level. And I feel very good about like impact that people can have in those kinds of roles. I think conversely there's there are definite problems in my mind for um uh for basically aligning yourself way too much with the frontier labs, too. So fundamentally I mean you're you have a huge amount of financial incentive to uh with these frontier labs. And by your own admission, the uh the AIs are going to like really change humanity and society in very dramatic ways. And here you are basically like building the technology and benefiting from it like it and being like very allied to it through financial means. Like this was the conundrum that was in at the heart of you know, how OpenAI was started in the beginning. Like this was the conundrum that we were trying to solve. Mhm. Um and so you know, that so it's kind of um It's still not resolved.
他们正在主动地把自己自动化掉,而这恰恰就是他们所有人都在努力做的事。是啊。我去转了一圈——其中一些研究员也会感受到那种精神焦虑,对吧?因为他们能感觉到它在起作用,对吧?于是他们就想,我也完蛋了。我确实花了一阵子在 OpenAI 各处转悠,我当时就说,你们意识到没有,如果我们成功了,我们就都失业了——这事就会变成我们其实是在给 Sam 之类的人构建自动化。是给我,还是给董事会,我也不太确定,但反正就是——他们就是在构建所有这些自动化,给董事会或者 CEO 之类的。然后我们全都失业,可能还能在边上贡献一点。所以从这个角度看,确实让人有点心神不宁。我可以问你 Noam 的那个问题吗?嗯。你知道,你本来可以去做那件事的,对吧?在某个前沿实验室里,用大规模算力、和一群同事一起做自动研究 [清嗓子]。为什么不呢?嗯,我之前在那儿待过一阵子,对吧?而且我后来还重新回去过。所以某种程度上我是同意的,我觉得这个问题可以从很多个角度来切。这个问题有点带点预设。我得说,我对人们能在前沿实验室之外做出的贡献和影响力感觉非常好——很显然不是在产业内部,而是在更偏生态系统层面的角色里。比如你的角色就更偏生态系统层面。我现在的角色某种程度上也更偏生态系统层面。我对人们在这类角色里能产生的影响感觉非常好。反过来说,在我看来,把自己跟前沿实验室绑定得太深,是有明确的问题的。从根本上说,你跟这些前沿实验室之间存在巨大的财务利益绑定。而按你自己的说法,AI 会以非常剧烈的方式真正改变人类和社会。而你呢,基本上就是在构建这项技术、并从中获益,通过财务手段跟它紧紧结盟。这恰恰就是当初 OpenAI 创立时核心处那个难题。这正是我们当时想要解决的那个难题。嗯。所以你知道,那个——它某种程度上——它至今仍未解决。
[45:50]
is still not like fully resolved. So that's number one. You're you're not a completely free agent and you can't actually like be part of that conversation in a fully autonomous um free way. Like if you're inside one of the frontier labs. Like there's some things that you can't say. Uh and conversely there are some things that the organization wants you to say. And you know, they're not going to twist your arm, but you feel the pressure of like what you should be saying, you know, cuz like obviously
至今仍然没有被完全解决。这是第一点。你并不是一个完全自由的行动者,你其实没办法以一种完全自主、完全自由的方式参与那场对话。如果你身处某个前沿实验室内部,有些话你就是不能说。反过来,有些话则是组织希望你去说的。他们不会硬掰你的胳膊逼你,但你能感受到那种压力——你应该说什么样的话,你懂的,因为很显然
[46:13]
[laughter]
[笑声]
[46:14]
otherwise it's like really awkward conversations, uh strange side eyes, like what are you doing, you know, like so you can't like really be an independent agent. And I I feel like a bit more a lot like aligned with humanity in a certain sense outside of the frontier lab because I don't I'm not subject to those pressures almost, right? And I can say whatever I want or Yeah, I would say in the frontier labs like um you can have like impact there of course as well. So but there's many researchers and maybe you're one of them, maybe your ideas are really good, etc. Maybe there's a lot of decision-making to do and you want to be in a position where you are in the room with those conversations when they come up. I do think that currently the stakes are like overall fairly low and so everything is kind of like nice. But ultimately in the end of the day like when the stakes are really high, etc. If you're an employee at an organization, I don't actually know how much sway you're going to have on your organization what it's going to do. Like fundamentally at the end of the day um uh it's uh you're not like really in charge. Like you're in the room and you're contributing ideas, but you're not like really in charge of that entity that you're that you're part of. So those are like some sources of misalignment, I think to some extent. I will say that like in one way I do agree a lot with that sentiment that um I do feel like in the like the labs for better or worse they're opaque and a lot of work is there. And they're kind of like at the edge of capability and what's possible. And they're working on what's coming down the line. And I think if you're outside of that frontier lab, your your judgment fundamentally will start to drift because you're not part of the you know, what's coming down the line. And so I feel like my judgment will inevitably start to drift as well. And I won't actually have an understanding of how these systems actually work under the hood. That's an opaque system. I won't have a a good understanding of how it's going to develop and etc. And so I do think that in that sense I agree and something I'm nervous about. I think it's worth basically being in touch with what's actually happening and actually being in a frontier lab. And if if some of the frontier labs would have me come for you know, some amount of time and do really good work for them and then maybe come and hang out.
否则就会出现非常尴尬的对话、奇怪的侧目而视,那种「你在干什么」的眼神,你懂的,所以你没办法真正做一个独立的行动者。我觉得在前沿实验室之外,某种意义上我感觉自己跟人类的立场对得更齐一些,因为我几乎不受那些压力的约束,对吧?我想说什么就能说什么。是啊,我得说,在前沿实验室里你当然也能产生影响。那里有很多研究员,也许你就是其中之一,也许你的点子真的很好,等等。也许有大量的决策要做,而你希望自己处在一个位置上——当那些对话发生时,你就在房间里。我确实觉得,眼下整体来说赌注还相当低,所以一切都还挺美好的。但归根结底,到了最后,当赌注真的非常高的时候,等等——如果你是一个组织的雇员,我其实并不知道你对自己的组织、对它将要做什么能有多大的左右力。从根本上说,到头来——你并不是真正的掌权者。你在房间里,你在贡献想法,但你并不是真正掌控你所属的那个实体的人。所以这些某种程度上就是一些立场不一致的来源吧。我得说,从某个角度看,我也很认同那种说法——我确实觉得,无论好坏,这些实验室是不透明的,大量的工作都在那里发生。它们某种意义上处在能力的边缘、处在「什么是可能的」的边缘。它们在做的是接下来要到来的东西。我觉得如果你身处那个前沿实验室之外,你的判断力从根本上会开始漂移,因为你不是「接下来要到来的东西」的一部分。所以我觉得我的判断力也将不可避免地开始漂移。我其实也不会真正理解这些系统在引擎盖底下究竟是怎么运作的。那是一个不透明的系统。我也不会很好地理解它将会怎么发展,等等。所以在这个意义上我确实是认同的,这也是我感到紧张的一点。我觉得保持跟实际正在发生的事情的接触、真正待在一个前沿实验室里,基本上是有价值的。如果某些前沿实验室愿意让我去待上一段时间、为他们做一些真正出色的工作,然后也许再出来一起玩玩。
[48:00]
looking for a job. This is super exciting. [laughter] Then I think that's maybe a good setup because I kind of feel like it's kind of um you know, maybe that's like one way Mhm. uh to to actually be connected to what's actually happening, but also not feel like you're necessarily fully controlled by Yeah. by those entities. So I think honestly in my mind like Noam can probably get do extremely good work at at OAI, but also I think his most impactful work could very well be outside of OpenAI. Noam, that's a call to be an independent researcher with auto [laughter] research. Yeah, there's many things to do on the outside and it's it's a and I think ultimately I think the ideal solution maybe is like yeah, going back and forth or um yeah, and I think fundamentally you can have a really amazing impact in both places. So very complicated I don't know. Like it's a very loaded question a little bit, but I mean I joined the frontier lab and I'm outside. And then maybe in the future I'll want to join again. And I think um uh that's kind of like how I look at it. One question related to what visibility to does the world or the AI ecosystem have into the frontier is like how how close open source is to the frontier. Mhm. Um and how sustainable that is. I I think Yeah. I think it is quite surprising. The entire sequence of events actually from like having a handful of Chinese models and global models and I think people are going to continue releasing here in the near term that are closer than much of the industry anticipated from a capability [clears throat] perspective.
找工作呢。这太令人兴奋了。[笑声] 那我觉得这也许是个不错的安排,因为我有点觉得这某种程度上——你知道——也许这是一种方式,嗯,去真正跟实际正在发生的事情保持连接,但同时又不会觉得你必然被那些实体完全控制。是啊。被那些实体。所以老实说,在我心里,我觉得 Noam 大概能在 OpenAI 做出极其出色的工作,但我也觉得他最有影响力的工作很可能就在 OpenAI 之外。Noam,这是在号召你做一个带着自动研究 [笑声] 的独立研究员。是啊,外面有很多事情可以做,这是一种——我觉得归根结底,理想的解决方案也许就是来回切换吧,或者,嗯,我觉得从根本上说,你在这两个地方都能产生真正了不起的影响。所以这事很复杂,我也说不准。这个问题确实有点带点预设,不过我的意思是,我加入过前沿实验室,现在又在外面。然后也许将来我又会想再加入一次。我觉得,嗯,我大概就是这么看待这件事的。还有一个相关的问题——世界、或者说 AI 生态系统对前沿有多少可见度,比如说开源离前沿有多近。嗯。以及这种状态有多可持续。我觉得——是啊,我觉得这相当令人惊讶。整个事件的演进序列其实——从一开始只有寥寥几个中国模型和全球模型,到现在——我觉得人们在近期还会继续发布一些模型,从能力 [清嗓子] 角度看,它们比业界很多人预期的要更接近前沿。
[49:26]
Yeah. Um I don't know if you're surprised by that, but you're a long-term contributor to open source. Like what's your prediction here? Yeah, so roughly speaking basically the the closed models are ahead, but like people are monitoring the number of months that sort of like open-source models are behind. Um And started with there's nothing and then it went to 18 months. Now it's
是啊。嗯,我不知道你对此是否感到惊讶,不过你是开源的长期贡献者。你对这件事的预测是什么?是啊,所以大致来说,基本上闭源模型是领先的,但人们一直在盯着开源模型大约落后多少个月这个数字。嗯,一开始是什么都没有,然后变成了 18 个月。现在是
[49:41]
Yeah, but then convergence, right? So then maybe they're behind by like, what is the latest? Maybe like 8 months, 6 months, 8 months kind of thing right now. Yeah, I'm a huge fan of open-source, obviously. So for example, in operating systems, you have like closed source, like, you know, Windows and Mac OS, these are large software projects, kind of like what LLMs are going to become, and there's Linux. Mhm. But Linux is very easy. Like, actually Linux is extremely successful project. It runs on the vast majority of computers. Like, last time I checked, was it like 60% or something like from Linux? Um and that's because there is a need in the industry to have a common open platform that everyone feels uh sort of safe using. I would say like the industry has always felt a demand for that kind of a project to exist. Mhm.
是啊,但接着就是收敛了,对吧?所以现在也许它们落后大概——最新的说法是多少来着?也许差不多 8 个月、6 个月、8 个月这个量级吧,现在。是啊,我显然是开源的超级粉丝。比如说在操作系统领域,你有闭源的,比如 Windows 和 Mac OS,这些都是大型软件项目,有点像 LLM 将来会变成的样子,然后还有 Linux。嗯。但 Linux 非常——其实 Linux 是一个极其成功的项目。它运行在绝大多数计算机上。我上次查的时候,是不是有大概 60% 之类的比例是跑在 Linux 上的?而这是因为产业里有一种需求,需要有一个共同的、开放的平台,让每个人用起来都觉得某种程度上是安全的。我得说,产业一直对这样一个项目的存在抱有需求。嗯。
[50:16]
And I think the same is true now. And that's why businesses actually want there's demand for this kind of a um a thing to exist. The big difference is that everything is capital uh there's a lot of capex that goes into this.
我觉得现在也是一样的道理。这就是为什么企业其实想要——存在着对这样一个东西存在的需求。最大的区别在于,这件事一切都很吃资本,里面有大量的资本支出(capex)。
[50:27]
Um so I think that's where things like fall apart a little bit, make it a bit harder to to compete in certain senses. Uh I I do think that the current models are very good. The other thing that I think is like really interesting is that for the vast majority of like consumer use cases and things like that, even like turn open-source models are actually quite good, I would say. And I think like if you go forward like more uh more years, it does seem to me like a huge amount of like simple use cases are going to be well covered and actually even run locally. Mhm. Um but there's going to be always like some demand for like frontier intelligence and that that can actually be extremely large uh piece of the pie. But it could be that the frontier the need for frontier intelligence is going to be like, you know, Nobel Prize kind of work. Mhm.
嗯,所以我觉得这就是事情有点崩盘的地方,让它在某些意义上更难去竞争。我确实觉得目前的这些模型非常好。另一件我觉得真的很有意思的事情是,对于绝大多数的消费级使用场景之类的东西来说,我得说,即便是当下的开源模型其实也已经相当不错了。我觉得如果你往后再走几年,在我看来,海量的简单使用场景都会被很好地覆盖到,甚至能在本地运行。嗯。但永远会有某种对前沿智能的需求,而那一块其实可能是极其庞大的一块蛋糕。不过也有可能,对前沿智能的需求会变成那种——你知道——诺贝尔奖级别的工作。嗯。
[51:05]
let's move Linux from C to Rust. It's going to be like bigger projects, you know, like scoped in that kind of a way, and there's going to be maybe more um and maybe that's where a lot of the frontier closed intelligence is where going to are going to be interacting with. And open-source kind of like going to eat through a lot of the more basic use cases or something like that. You know, at some point what is frontier today is going to be, you know, probably later this year what's frontier today in terms of what I'm using right now from the closed labs uh might be open-source and that's going to be doing a lot of work. So I kind of expect that this dynamic will actually basically continue. Like we'll have frontier labs that have closed um AIs that are kind of like these oracles, and then we'll have open-source kind of like behind with some amount of months. And I kind of expect that to uh to continue. And I actually think that's like a pretty pretty good setup uh overall. Um because I I'm a little bit hesitant of having um I don't actually think it's like structurally I think there's some systemic risk attached to just having intelligence that are closed and that's like that's it. Mhm. And I think that that's a, you know, centralization has a very poor track record in my view uh in in the past and has um
比如说把 Linux 从 C 迁移到 Rust。它会变成那种更大型的项目,你知道,按那种规模来界定的项目,可能还会有更多——也许那才是大量前沿闭源智能将要去打交道的地方。而开源某种程度上会逐步啃掉大量更基础的使用场景之类的。你知道,到某个时候,今天算是前沿的东西——你知道——大概在今年晚些时候,今天在我看来算前沿的、我现在正在用的这些来自闭源实验室的东西,也许就会变成开源的,而那些开源模型会承担大量的工作。所以我有点预期这个动态其实基本上会持续下去。就是说,我们会有前沿实验室,它们拥有闭源的、有点像神谕(oracle)一样的 AI,然后我们会有开源的、落后几个月这样的东西。我有点预期这个格局会持续下去。我其实觉得整体来说这是一个相当相当不错的格局。因为我对——我有点不太愿意看到那种——我其实不觉得,从结构上说,我觉得只拥有闭源的智能、就这样、没别的,是带有某种系统性风险的。嗯。我觉得那是一个——你知道——在我看来,中心化在过去有着非常糟糕的历史记录,而且还有——
[52:07]
You mean like in political or economic systems in in general.
你是指总体上在政治或经济系统里的那种中心化。
[52:10]
[laughter]
[笑声]
[52:12]
Exactly. I think there's like a lot of like pretty
没错。我觉得有很多挺……
[52:13]
an Eastern European. A lot of pretty bad precedents, so I want there to be a thing that is maybe not at the edge of capability because it's new and unexplored, etc. But I want there to be a thing that's behind and that uh is kind of like a common working space for intelligences that the entire industry has access to. Yeah, that seems to me like a pretty decent power balance for the industry. Yeah. I also think there's just like there are many problems to solve, right? Like if you keep advancing intelligence from the frontier, we can do new things and there are a lot of like very big problems for humanity, right? And so like it seems that that will continue to be a very expensive game. And so I want to like root for labs that are doing that because there are problems we cannot solve without continuing to advance the models in a very expensive way. And yet, as you point out, like if what we have today as frontier is open, that's a lot of capability, right? And and so I I I think, you know, the power of that or the democratization of that seems like
(像东欧那样的)。有很多挺糟糕的先例,所以我希望存在这么一个东西——它也许不在能力的最前沿,因为最前沿的东西是全新的、还没被探索过的,等等。但我希望有这么一个稍微落后一点的东西,它是整个行业都能用上的、智能体的某种公共工作空间。是的,在我看来这对整个行业来说是个挺不错的权力平衡。对。我还觉得,要解决的问题实在太多了,对吧?如果你能不断把智能从前沿往前推,我们就能做新的事情,而人类面临着很多非常重大的问题,对吧?所以看起来这会一直是个非常烧钱的游戏。因此我想为那些在做这件事的实验室加油,因为有些问题如果不以非常烧钱的方式继续推进模型,我们就根本解决不了。但同时,正如你指出的,如果我们今天的前沿是开放的,那就意味着大量的能力被释放了,对吧?所以我觉得,那种力量,或者说那种能力的民主化,看起来……
[53:04]
Yeah. very useful and also healthy.
对,非常有用,而且也很健康。
[53:06]
Yeah. I think basically by accident we're actually like in an okay spot.
对。我觉得基本上是误打误撞,我们其实正好处在一个还不错的位置。
[53:09]
An optimal. Yeah. [laughter] Yeah. Like by accident we we are it happened to be in a good spot in a certain sense. Mhm. Um Well, and and to some degree the the longer this endures, like this dynamic, um the the the healthier of a spot like the ecosystem might be in, right? Because you have more and more area under the curve.
一个最优的位置。对。[笑声] 是啊,纯属误打误撞,从某种意义上说我们正好落在了一个挺好的位置上。嗯。而且某种程度上,这种动态持续得越久,整个生态系统所处的位置可能就越健康,对吧?因为曲线下的面积会越来越大。
[53:25]
Mhm. And I will say that even on the closed side, I I almost feel like it's been like even further centralizing recently because I think a lot of the frontrunners are like not necessarily like the top tier. And so uh yeah, like in that sense I think it's um it's not super ideal. I would love there to be more more frontier labs because yeah, I'm like by default very suspicious of like um I want there to be more people in the room. I want I think like in machine learning ensembles always outperform any individual model. And so I want there to be ensembles of people thinking about all the hardest problems and I want there to be ensembles of people in the room when they um to be all well informed and to make those decisions, you know, so uh I don't want it to be like a closed doors with two people or three people. I feel like that's like not a good not a good future. I almost wish like there were more labs as long as they're short and I I I do think that open-source has a has a has a place to play. I hope it sticks around and I basically I it's currently slightly behind and it's actually kind of like a good thing. Okay, you worked on the precursor to generalized robotics autonomy um in cars, right? Uh a a lot has happened in the last couple months with robotics companies as well, like acceleration of really impressive generalization of environment, of tasks, like increasingly long horizon tasks, lots of money going into the space. Like, is it going to happen? Has anything in your view changed recently? Uh so like my view is kind of informed by what I saw in self-driving and I do feel like self-driving is the first robotics application. So probably what I saw is at the time, like 10 years ago, there were a large number of startups. And I kind of feel like um like most of them basically like didn't long-term make it. Um and what I saw is that like a lot of capital expenditure had to go in and a lot of time. And so um I think it's like I think robotics, because it's so difficult, is so messy, and requires a huge amount of capital investment, and a lot of like conviction. Um just it's like a big problem and I think atoms are really hard. So I kind of feel like they will lag be it will lag behind what's going to happen in digital space. And in digital space there's going to be a huge amount of unhobbling, uh basically like things that weren't super efficient becoming a lot more efficient by like a factor of a hundred.
嗯。我还想说,即便在闭源这一边,我都觉得最近反而更进一步中心化了,因为我觉得很多领跑者并不一定是真正的顶尖梯队。所以从这个意义上说,我觉得现状不算特别理想。我特别希望能有更多前沿实验室,因为说实话,我天生就对……我希望房间里有更多的人。我觉得在机器学习里,集成(ensemble)总是比任何单个模型表现更好。所以我希望有一群人组成的集成去思考所有最难的问题,我希望当他们做那些决策时,房间里有一群人——大家都消息灵通,一起做出那些决定。所以我不希望它变成两三个人关起门来做决定的局面。我觉得那不是个好的未来。我几乎是希望有更多的实验室,只要它们还紧跟前沿就行。而且我确实认为开源有它该扮演的角色。我希望它能一直存在下去。基本上,开源现在稍微落后一点点,而这其实反倒是件好事。好,你曾经做过通用机器人自主性的前身工作——就是在汽车领域,对吧?过去几个月机器人公司也发生了很多事,比如对环境、对任务的泛化能力出现了非常惊人的加速,任务的时间跨度越来越长,大量资金涌入这个领域。那么,这件事会成吗?在你看来最近有什么变化吗?我的看法很大程度上是被我在自动驾驶里看到的东西塑造的,我确实觉得自动驾驶是第一个机器人应用。所以我当时看到的——大概十年前——有非常多的创业公司。我感觉它们当中大多数从长期来看基本都没能挺过来。我看到的是,必须投入大量的资本开支,还有大量的时间。所以我觉得机器人这件事,因为它太难了、太混乱了,需要巨额的资本投入和大量的信念。它就是个大难题,我觉得原子(实体世界)真的非常难。所以我感觉机器人会滞后于数字空间里将要发生的一切。而在数字空间里,会出现大量的「解绑」(unhobbling)——基本上就是那些原本效率不高的东西,效率会一下子提升一百倍。
[55:25]
Mhm. Because bits are so much easier. And so I think currently in terms of what's going to change and like where the activity is, I kind of feel like digital space is going to like change a huge amount. And then the physical space will lag behind. And what I find very interesting is like this interface in between them as well. Because I think in this like if you we do have more agents acting on behalf of humans and more agents kind of like talking to each other and and doing tasks and participating in kind of economy of agents, etc. Um you're going to run out of things that you're going to do purely in the digital space. At some point you have to go to the universe and you have to ask it questions. Um you have to run an experiment and see what the universe tells you to get back to learn something. And so we currently have a huge amount of like digital work uh because there's an overhang in how much we collectively thought about what already is digital. So we just didn't have enough thinking cycles among the humans to think about all the information that is already digital and already uploaded. Um and so we're going to start running out of stuff that is actually like um already up uploaded. Uh so you're going to at some point read all the papers and process them and have some ideas about what to try, but um yeah, we're just going to uh I don't actually know how much you can like get intelligence that's like fully closed off and was just information that's available in the you know. And so I think what's going to happen is first there's going to be a huge amount of unhobbling and I think there's a huge amount of work there. Then actually it's going to move to like the interfaces between physical and digital. So I and that's like sensors of like seeing the world and actuators of like doing something to the world.
嗯。因为比特(数字世界)要容易得多。所以我觉得,就当下来说,论将要发生变化的地方、论活力所在,我感觉数字空间会发生巨大的变化,而物理空间会滞后。我觉得特别有意思的是它们之间的这个接口。因为我想,如果我们确实会有越来越多代表人类行动的智能体、越来越多互相交流、执行任务、参与到「智能体经济」中的智能体,等等,那么你会发现纯粹在数字空间里能做的事情迟早会做完。到了某个点,你必须走向真实宇宙,向它提问。你必须去做实验,看宇宙给你什么反馈,才能学到东西。我们现在之所以有海量的数字化工作可做,是因为我们对「已经数字化的东西」的集体思考存在一个积压(overhang)。也就是说,人类的思考算力还不够,没能把那些已经数字化、已经上传的所有信息都想透。所以我们迟早会把那些已经上传的东西用完。你迟早会把所有论文读完、处理完,对该尝试什么有些想法,但是……总之我们会……我其实不知道,单靠那种完全封闭起来的、只用已有信息的方式,到底能造出多强的智能。所以我觉得接下来会发生的是:先有大量的「解绑」,我认为那里有海量的工作要做。然后真正会转移到物理与数字之间的那些接口上。也就是说,「看见世界」的传感器,和「对世界做点什么」的执行器。
[56:48]
Mhm. So I think a lot of interesting companies will actually come from that interface of like can we feed the superintelligence in a certain sense uh data and can we actually like take data out and manipulate the physical world um per its bidding if you want to like anthropomorphize the whole thing, right? And then the the physical world actually I almost feel like the the total addressable market, etc. in terms of like the amount of work and so on is is massive, possibly even much larger maybe what can happen in digital space. So actually think it's like a much bigger opportunity as well. But um I do feel like it's a huge amount of work and and in my in my mind the atoms are just like a a million times harder. So um so it will lag behind, but it's also I think a little bit of a bigger market. So it's kind of like uh yeah, I think the opportunity is kind of like follow that kind of trajectory. So right now is digital is like my main interest. Then interfaces will be like after that and then maybe like some of the physical things um like their time will come and they'll be huge when they do come. Well, it's it's it's an interesting framework for it, too, because uh certain things, not the things I'm working on right now, but certain things are much easier even in the world of atoms.
嗯。所以我觉得会有很多有意思的公司其实会从那个接口里冒出来——比如说,从某种意义上讲,我们能不能给 superintelligence「喂」数据,我们能不能真的把数据取出来、按照它的「指令」去操纵物理世界——如果你想把整件事拟人化的话,对吧。然后说到物理世界,我几乎觉得,论总可触达市场(TAM)之类的、论需要做的工作量等等,那是个巨大的市场,可能甚至比数字空间里能发生的还要大得多。所以我其实觉得它也是个大得多的机会。但是,我确实觉得那需要海量的工作,在我心里,原子就是比比特难一百万倍。所以它会滞后,但它同时也是个稍微更大一点的市场。所以这件事大概就是……对,我觉得机会大概会沿着这样一条轨迹走。所以现在我主要的兴趣在数字。然后接下来是接口。再之后也许是一些物理层面的东西——它们的时机会到来,而当它们到来时,会非常巨大。这也是个挺有意思的框架,因为有些事情——不是我现在在做的这些——但有些事情即便在原子的世界里其实也容易得多。
[57:51]
Mhm. Right? Like if you just think about like read and write to the physical world, like read, like sensors, cameras, like there's a lot of existing hardware and you can imagine like enriching agent capabilities or capturing a lot of new data if you just clever about it and like you don't necessarily have to invest a lot to like get something valuable.
嗯。对吧?比如你想想看,对物理世界的「读」和「写」——「读」就是传感器、摄像头之类,已经有大量现成的硬件。你可以想象,只要你足够聪明,就能丰富智能体的能力,或者采集大量新数据,而你不一定非得投入很多就能拿到有价值的东西。
[58:10]
Yeah. Right. Yeah. So like examples of this that I saw for example are, you know, um a friend of mine, Liam, is running is a CEO of Periodic. I visited them last week. Yeah. So it was just on top of mind. Like they're trying to do auto research for materials science. Mhm. Um and so in that case it's like the sensors to the intelligence are actually like pretty expensive lab equipment. And the same is true in biology. I think a lot of people are very interested in engineering biology and, you know, the sensors will be more than just like video cameras. Does that make sense? And then the other thing I was I saw for example is companies that are trying to have um like you basically pay people for training data. Yeah. Yeah. Yeah. Yeah.
对。是的。所以这方面我看到的例子,比如说,我有个朋友叫 Liam,他是 Periodic 的 CEO,我上周去他们那儿拜访过。对,所以这事正好在我脑子里。他们想做的是材料科学领域的自动研究(auto research)。嗯。在那个场景里,通向智能的那个「传感器」其实是相当昂贵的实验室设备。生物领域也是一样。我觉得很多人对工程化生物学非常感兴趣,而那里的「传感器」就不只是视频摄像头那么简单了。明白我的意思吗?另外我还看到的例子是,有些公司在尝试……基本上就是你花钱请人来生产训练数据。对,对,对。
[58:42]
To feed the Yeah.
用来喂养那个……对。
[58:42]
programmatically.
以程序化的方式。
[58:43]
Yeah. To feed to feed the Borg. Uh um and so like these are all examples of like sensors in a certain sense. So they take many diverse shapes and forms if that makes sense. Mhm. Yeah, so I'm looking forward to the point where I can ask for a task in the physical world and I can put a price on it and just tell the agent like, you know, you figure out how to do it. Go get the data.
对。用来喂养 Borg(博格集合体)。所以从某种意义上说,这些全都是「传感器」的例子。它们会以五花八门的形态出现,如果你懂我意思的话。嗯。是啊,所以我很期待这样一天的到来:我能在物理世界里下达一个任务,给它标个价,然后直接告诉智能体——你自己想办法搞定,去把数据弄回来。
[59:02]
I'm actually kind of surprised we don't have enough like information markets. Mhm. Like if for example if Polymarket or other betting markets or even stocks, etc. If they have so much autonomous activity and rising amount of activity, Mhm. like um why should like for example if Iran was just happening now, like how come there isn't a process where like taking a photo or video from somewhere in Tehran should cost like 10 bucks? Like someone should be able to pay for that, you know, like and that's an example of like feeding the intelligence. There's not going to be a human looking at it, it's going to be like agents who are trying to guess the betting games and stock markets and so on. Mhm. So I kind of feel like the agentic web is still like fairly new, but there's no like mechanisms for this, but this is an example of what I I think might happen. Uh there's a good book that maybe is inspiring called Daemon. Mhm. You potentially read it. In Daemon, the intelligence um ends up like puppeteering almost a little bit like humanity in a certain sense, you know? And so, humans are kind of like it's actuators, but humans are also like its sensors. Um and so, I think like collectively like society will kind of like reshape in a certain way in uh to to serve that kind of a that will kind of like end up happening collectively across the industry. Where yeah, there's just a lot more automation and it has certain needs and kind of humans will be serving those needs of that of that machine, not necessarily like to each other.
我其实有点惊讶我们居然还没有足够多的「信息市场」。嗯。比如说,如果像 Polymarket 或者其他博彩市场、甚至股市等等——如果它们有那么多自动化的活动,而且活动量还在上升——嗯——那为什么……比如说,如果伊朗那档子事现在正在发生,怎么就没有一种机制,让「从德黑兰某处拍张照片或视频」值大概 10 美元呢?应该有人愿意为这个付费才对,对吧?而这就是「喂养智能」的一个例子。不会有人类去看那张照片,看它的会是那些试图押注博彩游戏、股市之类的智能体。嗯。所以我感觉「智能体的网络」(agentic web)还相当新,目前还没有这类机制,但这就是我觉得可能会发生的事情的一个例子。有一本挺好的书也许能给人启发,叫《Daemon》(守护程序)。嗯。你也许读过。在《Daemon》里,那个智能最终从某种意义上几乎是在「操纵」整个人类,懂吗?所以人类有点像是它的执行器,但人类同时也是它的传感器。所以我觉得,整个社会会以某种方式集体地重新塑形,去服务于那种……这种事会在整个行业里集体地发生。也就是,会有多得多的自动化,而它有某些需求,人类某种程度上会去服务那台机器的需求,而不一定是去服务彼此。
[1:00:12]
Well, we were um on this very specific point of uh like missing pieces of training data. We needed um we needed something like auto research, right? Like we we need the training cycle or the SFTP piece to be uh far more mechanized. Mhm. For for which part?
对,我们刚才正说到一个非常具体的点——就是训练数据里缺失的那些拼图。我们需要某种类似「自动研究」的东西,对吧?我们需要让训练循环、或者说 SFT(监督微调)那一环变得机械化得多。嗯。是为了哪一部分呢?
[1:00:28]
In order to make the uh collection like to in order to take the human out of the loop to ask for a task that is just like improve my model quality with new data, right? Uh yes. Does that make sense to you? Like we um if you can't have the model do the training runs by itself, then your ability to do this as a like closed loop task with uh by pricing data is um more challenged. Yes, yes, 100%. Yeah. But now you do.
是为了让数据采集……为了把人从循环里拿掉,让你能下达这样一个任务:「用新数据来提升我的模型质量」就行了,对吧?嗯,对。你明白我的意思吗?就是说,如果你没法让模型自己去跑训练,那你把这件事做成一个「用价格去买数据」的闭环任务的能力就会受到更多限制。对,对,百分之百。是的。但现在你可以做到了。
[1:00:57]
The thing is for LLM training, it actually is like very easily it like really fits the paradigm. Mhm. Um so, you'd actually expect
问题是,对于 LLM 训练来说,它其实非常容易……它真的非常契合这个范式。嗯。所以你其实会预期……
[1:01:04]
metric. Yeah, like LLM training actually fits the paradigm really well, really easily. Like all the optimization of all the code and so, it runs faster. And then you also have like metrics that you can optimize against. I do think that if you had an autonomous loop over those metrics, there's going to be a lot of like good herding going on where the system will like overfit to those metrics. And so, um but then you can use the system to devise more metrics and you just have a really good coverage. So, it's kind of hard to tell, but um in a certain sense it's like a pretty pretty good fit. I want to talk about a little uh tiny side project you have before we end. Um tell me about the micro GPT arts. Oh, yeah. Okay, so micro GPT. So, I have this like running obsession of like maybe a decade or two of just like simplifying and boiling down the uh basically LLMs uh to like their bare essence. And I've had a number of projects along these lines. So, like nano GPT and um make more and uh micro GPT micro grad etc. So, I feel like micro GPT is now the state of the art of me trying to like just boil it down to just the essence. Because the thing is like training neural nets and LLMs specifically um is a huge amount of code, but all of that code is actually complexity from efficiency. It's just because you need it to go fast. If you don't need it to go fast and you just care about the algorithm, then that algorithm actually is uh 200 lines of Python, very simple to read. And this includes comments and everything. Um because you just have like uh your data set which is a text um and you need your neural network architecture which is like 50 lines. You need to do your forward pass and then you have to do your backward pass to calculate the gradients. And so, an auto grad engine uh to calculate the gradients like 100 lines. And then you need an optimizer and Adam for example, uh which is a very state of the art optimizer is like again 10 lines, really. And so, putting everything together in the training loop is like yeah, 200 lines. And what's interesting to me like normally before like maybe a year ago or more, if I had come up with micro GPT, I would be tempted to basically explain to people. Like I have a video like stepping through it or something like that. Uh and I actually tried to make that video a little bit. And I tried to make like a little guide to it and so on. But I kind of realized that this is is not really is not really adding too much because people cuz it's already so simple that it's 200 lines that anyone could ask their agent to explain it in various ways. And the agents like I'm not explaining to people anymore. I'm explaining it to agents. If you can explain it to agents, then agents can be the router and they can actually target it to the human in their language uh with infinite uh you know, patience and uh just at their capability and so on. Right. If I don't understand um this particular function, I can ask the agent to explain it to me like three different ways and I'm not going to get that from you. Exactly. And so, I kind of feel like, you know, what is education? Like it used to be guides, it used to be lectures, it used to be this thing, but now I feel like now more I'm explaining things to agents and maybe I'm coming up with skills uh where like um uh so, basically skill is just a way to instruct the agent how to teach the thing. So, maybe I could have a skill for micro GPT of the progression I imagine the agent should take you through if you're interested in understanding the code base. And it's just like hints to the model to like uh first start off with this and then with that. And so, I could just script the curriculum a little bit as a skill. Uh so, uh so, I I don't feel like um yeah, I feel like there's going to be less of like explaining things directly to people and it's going to be more of just like does the agent get it? And if the agent gets it, they'll do the explanation. And we're not fully there yet because they I still can I still think I can probably explain things a little bit better than the agents, but I still feel like the models are improving so rapidly that um I feel like it's a losing battle to some to some extent. Um and so, I think education is going to be kind of like reshuffled by this quite substantially uh where it's the end of like teaching each other things a little bit like if I have a um library for example of code or something like that. It used to be that you have documentation for other people who are going to use your library, but like you shouldn't do that anymore. Like you should have instead of HTML documents for humans, you have markdown documents for agents. Cuz if agents get it, then they can just explain all the different parts of it. So, it's this redirection through agents, you know? Um and that's why. So, I think we're going to see a lot more of that playing out. Well, we'll see if the great teachers know like to develop intuition for how to explain things to agents differently.
(指标)。对,LLM 训练其实非常好、非常容易地契合这个范式。比如所有代码的各种优化,它跑得更快了。然后你还有可以去优化的指标。我确实认为,如果你在那些指标上套一个自主循环,会出现大量的「钻空子」(herding),系统会过拟合到那些指标上。但接下来你可以用这个系统去设计出更多的指标,于是你就能有非常好的覆盖度。所以很难说,但从某种意义上讲,这个范式契合得相当、相当好。在结束之前,我想聊一个你做的小小的边角项目。跟我说说 micro GPT 这个事吧。哦,对。好,micro GPT。我有一个持续了大概十年、二十年的执念,就是不断地简化、把 LLM 提炼到它最本质的东西。沿着这个方向我做过好几个项目,比如 nano GPT、makemore、micro GPT、micrograd,等等。所以我觉得 micro GPT 现在算是我把这件事提炼到只剩本质的最新成果。因为问题是,训练神经网络、尤其是训练 LLM,需要海量的代码,但所有那些代码其实都是为了「效率」而产生的复杂性。它们存在只是因为你需要它跑得快。如果你不需要它跑得快,你只关心算法本身,那这个算法其实只有 200 行 Python,非常好读。而且这还包括了注释和所有东西。因为你只需要一个数据集——就是一段文本——然后你需要你的神经网络架构,那大概 50 行。你需要做前向传播,然后你得做反向传播来计算梯度。所以一个用来算梯度的 autograd 引擎大概 100 行。然后你需要一个优化器,比如 Adam——它是个非常前沿的优化器——其实也就 10 行,真的。所以把所有东西拼进训练循环里,对,大概 200 行。让我觉得有意思的是:通常在……大概一年多以前,如果我搞出了 micro GPT,我会很想去给大家讲解它,比如做个视频一步步过一遍之类的。我其实也确实试着做了一点那样的视频,也试着写了个小指南之类的。但我后来意识到,这其实没增加多少价值,因为它已经简单到只有 200 行了,任何人都可以让他们的智能体用各种方式给他们解释。我现在不是在给人讲解了,我是在给智能体讲解。如果你能把它讲给智能体听,那智能体就可以当那个「路由器」,它能用人类自己的语言、以无限的耐心、按照人类的能力水平把它讲给人听。对吧。如果我不懂某个特定的函数,我可以让智能体用三种不同的方式给我讲,而这是我从你这儿得不到的。没错。所以我有点觉得,教育到底是什么?它过去是指南,是讲座,是这样那样的东西,但现在我感觉,我越来越多地是在把东西讲给智能体听,也许我在做的是「技能」(skill)——基本上,技能就是一种指示智能体如何去教某样东西的方式。所以也许我可以为 micro GPT 做一个技能,里面写着我设想智能体该带你走过的学习路径——如果你有兴趣理解这个代码库的话。它就是给模型的一些提示:先从这个开始,然后再讲那个。所以我可以把整个课程稍微编排成一个技能脚本。所以,我觉得,对,我感觉以后「直接给人讲解东西」会越来越少,而会越来越多地变成:智能体懂不懂?如果智能体懂了,它就会去做讲解。我们还没完全到那一步,因为我还是觉得我大概能比智能体讲得稍微好一点,但我还是觉得模型进步得太快了,以至于在某种程度上,这是一场注定要输的仗。所以我觉得教育会被这件事相当大幅地重新洗牌——会终结「人教人」这种东西。比如说,如果我有一个代码库之类的东西。过去你会为将要使用你这个库的其他人写文档,但现在你不该再那么做了。你不该再为人类写 HTML 文档,而应该为智能体写 markdown 文档。因为如果智能体懂了,它就能把库的各个部分都讲清楚。所以这是一种经由智能体的「转接」,懂吗?这就是原因。所以我觉得我们会看到越来越多这样的情况上演。那我们就看看那些真正出色的老师,会不会培养出一种「如何用不同方式给智能体讲解」的直觉吧。
[1:05:05]
ultimately, so for example, micro GPT, like I asked I tried to get an agent to write micro GPT. So, I told it like try to boil down the simplest things. Like try to boil down my um neural network training to the simplest thing and it can't do it. Like micro GPT is like my is it's like my end of my obsession. It's the 200 lines. I thought about this for a long time. I was obsessed about this for a long time. This is this is the solution. Trust me, it can't get simpler. And this is this is my value add. Everything else like agent gets it. It just can't come up with it, but it totally gets it and understands why it's done in a certain way etc. Uh so, like my contribution is kind of like these few bits, but everything else in terms of like the education that goes on after that is like not my domain anymore. So, maybe yeah, it's like education kind of changes in those ways where you kind of have to infuse the few bits that you feel strongly about the curriculum or the the best the better way of explaining it or something like that. The things that agents can't do is your job now. The things that agents can do, they can probably do better than you or like very soon. And so, you should um be strategic about what you're actually spending time on. Well, we appreciate the few bits. Thank you, Andre. Okay. Find us on Twitter at No Priors Pod.
归根结底,比如说 micro GPT,我试过让一个智能体来写 micro GPT。我跟它说,试着把最简单的东西提炼出来,把我的神经网络训练提炼到最简的形态,但它做不到。micro GPT 算是我执念的终点,就是那 200 行。我想这件事想了很久,我为这件事痴迷了很久。这就是那个解。相信我,它没法再更简了。这就是我增加的价值。其他所有东西,智能体都能搞懂。它只是想不出来这个解,但它完全能领会、能理解为什么要用某种特定方式去做,等等。所以我的贡献大概就是这么几个比特(bits),但在那之后所有关于「教育」的事情,就已经不再是我的领域了。所以也许,对,教育会以这样的方式发生改变——你得把那么几个你强烈认同的比特注入进去,比如课程的设计,或者某种最好的、更好的讲解方式之类的。智能体做不到的事,现在就是你的工作。智能体能做到的事,它们大概能比你做得更好,或者很快就会。所以你应该有策略地去想,你到底该把时间花在什么上面。好的,我们很感激这几个比特。谢谢你,Andrej。好。在 Twitter 上找我们,账号是 @NoPriorsPod。
[1:06:15]
[music]
[音乐]
[1:06:15]
Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. [music] That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.
如果你想看到我们的脸,就订阅我们的 YouTube 频道。在 Apple Podcasts、Spotify,或者任何你常用的平台上关注本节目。[音乐] 这样你每周都能收到一期新节目。还可以注册我们的邮件列表,或者在 no-priors.com 上找到每一期的文字稿。